Can AI Detectors Be Wrong? What Our Script Tests Show

By the AID Editorial Team · September 17, 2026

Can an AI detector be wrong? In AID’s initial tests, one detector missed a known AI-written script and another gave the same script an ambiguous explanation. These concrete results show why a confident-looking score needs context.

The AID answer: Use a detector result as one signal to investigate. It does not replace evidence of how a document was written. An uncertain result is still uncertain, and a missed AI sample does not become human-authored because a tool labels it that way.

What happened to our AI-written script?

Our eight-document protocol includes four human-written documents and four documents whose submitted wording was generated entirely by AI. One AI sample, AI03, was a script produced from detailed human instructions. The prose was not rewritten by a human before submission. Human direction matters to the creative process, but it did not change the known origin of this sample’s wording.

In our September 4, 2026 GPTZero evaluation, that script was classified as Human: the report showed 14% AI, 15% Mixed and 71% Human. It was the one incorrect classification in an otherwise seven-correct-result evaluation.

In our September 15, 2026 Winston AI evaluation, the script received a 39% Human Score with an ambiguous explanation. We recorded that ambiguity and awarded no correct-classification credit. We did not convert it into a confident human verdict or claim it was a clear AI detection.

A false negative and an ambiguous result differ

A false negative occurs when a known AI sample is classified as human. An ambiguous completed scan leaves the classification unresolved. Both can limit a detector’s usefulness, but they should be recorded separately. Otherwise a review can hide uncertainty by forcing every scan into a simple success-or-failure story.

Our published methodology gives ambiguous results no success credit. Incomplete scans are tracked separately as operational failures. This keeps the interpretation tied to what the tool actually returned.

What about false positives?

A false positive is the opposite error: human writing is classified as AI. All four human documents in each of our initial GPTZero and Winston evaluations were recognized as human. That is reassuring for these specific samples. Four human documents per product cannot establish how either tool treats all authors, genres, languages or editing workflows.

We therefore report no observed human false positives in these small tests. We do not promise that false positives never occur. An overall AID Score of 4.6/5 is an editorial evaluation across four equally weighted categories, not a claim of 92% classification accuracy.

Read the score before drawing a conclusion

Check what the number measures. Winston reports a Human Score; GPTZero displays separate AI, Mixed and Human probabilities. A 39% Human Score is not automatically a verified 61% AI probability. Preserve the tool’s original labels, explanation and report instead of inventing a conversion.

Also record the plan and date. Our Winston scans used paid Essential access. Seven GPTZero scans used Premium and one pilot used free access. These are dated evaluations of particular configurations, not evidence that every free plan performs identically.

A practical review process

Save the exact submitted text and detector report. Compare the result with drafts, source notes and revision history. If an important authorship question remains, discuss the writing process with the author and examine the relevant passages. Keep unresolved findings labeled as unresolved.

Running more tools may give additional signals, but agreement alone does not prove authorship. Our small evaluation set cannot establish that tools are independent of one another or justify treating their votes as a reliable probability.

For a no-cost starting point, read our Best Free AI Detectors guide. For the full evidence, use the individual reviews and AI Detector Directory. The strongest review explains both what the tool got right and where its answer became unreliable.

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top