Verification field guides

Two AI detectors disagree. What should you check next?

A disagreement is a reason to inspect the inputs and evidence. It is not a reason to average unrelated scores into a new certainty percentage.

By isGenAI · Published Reviewed

Confirm the inputs first

Were both tools given the same original, or did one receive a resized screenshot? Record dimensions, file type and the source of each copy. If you kept a hash, compare it. A mismatch means you are reviewing two assets, even if they look alike.

Also record the tool, date and the displayed explanation. A label without its scope is difficult to interpret later. Avoid reporting an unavailable result as a negative finding.

Sort evidence by the question it answers

Saved generation settings describe records present in a file. A supported credential provides an inspectable signed assertion. A model score estimates something from its input. A source statement describes what a publisher or creator says. Keep these categories visible instead of treating them as interchangeable votes.

  1. Preserve both complete results and identify their inputs.
  2. Inspect the original for saved records where available.
  3. Trace the source and compare the claim with the creator’s explanation.
  4. State the remaining uncertainty and the evidence that would resolve it.

Worked example: 90% accuracy does not mean every flag is 90% certain

This is a hypothetical arithmetic example, not an isGenAI evaluation or a claim about its scores. Assume 1,000 images with known, mutually exclusive labels: 100 AI-generated and 900 human-made. At a fixed decision threshold, suppose a detector flags exactly 90% of the AI images and correctly clears exactly 90% of the human images. Every image receives a decision; there are no abstentions or errors in running the check.

The counts are 90 correctly flagged AI images, 10 missed AI images, 90 falsely flagged human images and 810 correctly cleared human images. Overall accuracy is (90 + 810) / 1,000 = 90%. But only 90 of the 180 flagged images are AI-generated: precision among flags is 90 / (90 + 90) = 50%. Accuracy across the evaluation set and correctness among positive results have different denominators.

Now change only the assumed mix to 500 AI-generated and 500 human-made images, while holding both class-specific rates at 90%. There are 450 true flags, 50 false flags, 50 misses and 450 correct clearances. Accuracy remains 90%, but precision becomes 450 / (450 + 50) = 90%. Prevalence—the share of AI images in the population—changed from 10% to 50%. Real-world rates may also change with the input population; this example assumes they do not.

A displayed confidence score such as 0.90 is a separate quantity. For a well-calibrated classifier in an appropriate evaluation population, roughly 90% of comparable cases assigned scores near 0.90 should belong to the positive class. That requires calibration evidence, not just an accuracy figure. Neither of the two count examples supplies that evidence, and this guide makes no claim that isGenAI provides calibrated authorship probabilities.

Nor does a score measure the fraction of an image that was edited. In another hypothetical example, an independently known edit mask covering 20% of the pixels would describe an area measurement, not 20% probability of AI origin. A signed provenance record describing an editing action is another kind of evidence: inspect the recorded claim and its validation, rather than converting it into a confidence score or a percentage of generated pixels.

Sources: National Institute of Standards and Technology: NIST AI 100-4, Appendix E: evaluation metricsscikit-learn documentation: Probability calibrationCoalition for Content Provenance and Authenticity: C2PA and Content Credentials Explainer, version 2.2

Choose a proportionate conclusion

For a casual unexplained image, “origin unresolved” may be enough. For an accusation against a person, a detector disagreement is not adequate evidence of wrongdoing. Seek independent corroboration before making a consequential claim.

You can practice reading different evidence types using isGenAI’s controlled examples. Those examples show known file operations; they do not establish a universal accuracy rate for every image circulating online.

Sources and dates

Sources checked on 27 September 2026. Linked reporting and official statements establish the attributed facts; the review steps are isGenAI’s analysis.

Put the next question to the file.

Choose the supported tool that fits your question. Keep the original source and save the output with your notes. Frame extraction, file-record inspection and still-image checks answer different questions; none alone authenticates a complete video or its audio.

Our methodology explains supported checks and result limits. Found a factual error or a newer primary source? Use our corrections process and include this article’s URL.

Keep following the evidence