ARTIFICIAL INTEL DETECTOR

How the AID Score works.

Our transparent methodology for evaluating accuracy, false-positive protection, ease of use, and value.

TRANSPARENT SCORING

What the AID Score measures

Methodology version 1 · Updated September 17, 2026.
The AID Score combines four equally weighted ratings on a 1–5 scale: Accuracy, False-Positive Protection, Ease of Use, and Value. Each contributes 25%. We average the category ratings and round the overall result to one decimal. This is guidance based on the tested workflow and corpus, not an estimate of universal detection accuracy.

Calculation example: GPTZero’s initial ratings are 4.5, 5.0, 4.5, and 4.5. Their average is 4.625, displayed as 4.6/5.

25%

Accuracy

Accuracy measures clear, correct document classifications across the human and AI samples. We calculate 1 + (4 × the observed success rate), then round to the nearest half-star, with ties rounded up. Seven clear-correct results from eight documents gives 4.5/5. Completed ambiguous results receive no correct-classification credit under our conservative treatment; the review preserves the raw score and explains that decision. Failed or incomplete scans are not classifications and are reported separately.

25%

False-Positive Protection

False-Positive Protection measures the proportion of human documents correctly left unflagged. We use 1 + (4 × that success rate), rounded to the nearest half-star. Four of four gives 5.0/5 for this corpus; it does not establish a universal zero false-positive rate.

25%

Ease of Use

Ease of Use is an editorial rating in half-star steps, based on the workflow, clarity, reporting, limits, and effort we observed. A 4.5/5 means excellent with minor limitations. GPTZero lost half a star for a temporary history-status inconsistency. Larger deductions require material failures, with reasons documented consistently across products.

25%

Value

Value is an editorial rating in half-star steps, based on verified price, usable capacity, and relevant features. We assess the paid plan separately from free-plan restrictions. GPTZero earned 4.5/5 for useful paid capacity at the observed price; broader sustained-use and comparative checks are needed before assigning the maximum. A 4.5 rating is not a measured 90% success rate.

CONTROLLED TEST CORPUS

Known authorship—not assumptions

Our first benchmark uses eight long-form samples whose origin is known. Master copies are retained without detector-driven rewriting so every product receives the same evidence.

Four human-written samples

Four complete documents, 1,442–3,937 words, written by the AID editor-in-chief: book excerpts from 2021 and 2019, a doctoral-thesis excerpt from 2023, and a YouTube script from 2024. All come from one author, which limits generalization.

Four AI-written samples

Four complete documents with AI-generated wording, 1,555–2,603 words: two articles and a long-form book sample from 2026, plus a Claude-written script from 2025 created with extensive human direction. Ground truth follows who generated the submitted wording.

REPEATABLE PROTOCOL

How every detector is tested

01 — Lock the samples

The same complete master texts are used for every detector, without intentional rewriting, shortening, or tuning to affect a result. Input methods are documented: file upload and full plain-text submission can preserve layout differently. We record interface normalization and word-count differences rather than assuming that formatting and counts remain identical.

02 — Standardize conditions

Tests are run using the documented product tier, input method, date, and available settings. Material differences are recorded.

03 — Capture raw outputs

We preserve the detector's result, percentages or labels, highlighted passages, warnings, and any limitations shown during the test.

04 — Record both error types

We track missed AI writing, falsely flagged human writing, and ambiguous outputs separately. An ambiguous result is not automatically a false negative, even though it receives no success credit in our conservative accuracy calculation. A single combined accuracy number can hide these distinctions.

05 — Verify usability and price

We document workflow, reporting, word limits, account requirements, plan restrictions, and pricing as observed—not as remembered.

06 — Date and revisit

Every review shows when testing and pricing were last checked. Material product changes trigger a retest or a visible qualification.

RESPONSIBLE INTERPRETATION

Detection is a signal, not proof

No detector result, including a high probability or confident label, should by itself establish authorship, cheating, misconduct, or intent. High-impact decisions require human review, context, and corroborating evidence.

VERSION CONTROL

Results can change

Models, thresholds, interfaces, prices, and product policies evolve. Our scores describe the tested version under recorded conditions. Initial scores identify the corpus size and limitations; retesting may change them. A maximum rating in a small corpus is not proof of perfect performance across other authors, languages, or writing styles.

THE INITIAL EIGHT-DOCUMENT CORPUS

What is in our controlled test set

The first AID benchmark contains four verified human-written documents and four verified AI-written documents. We use the complete master files, not detector-selected excerpts. Descriptions are intentionally broad enough to protect private source material while still showing differences in genre, length, and production method.

H01 — Human long-form narrative

3,017 words. An original human-written long-form manuscript sample. It tests whether a detector can recognize sustained personal voice and narrative development without falsely labeling it as AI.

H02 — Human nonfiction book material

3,937 words. Original human-written material about mindset and professional experience. It includes personal perspective, structured explanation, and long-form instructional prose.

H03 — Human doctoral writing

2,557 words. Original human-written doctorate-level academic material. It represents formal research-oriented prose, where standardized language can sometimes resemble machine-generated text.

H04 — Human video-script writing

1,442 words. An original human-written YouTube script. It adds conversational pacing, audience address, and spoken-language structure to the human baseline.

AI01 — AI-written professional article

2,603 words. A complete AI-written article created for a professional audience. It tests polished explanatory prose in a conventional article format.

AI02 — AI-written long-form book material

2,557 words. Complete AI-written long-form material about AI and search. It represents structured, information-dense book-style writing.

AI03 — AI-written script with extensive human direction

1,789 words, 2025. Claude generated all submitted prose, so the ground truth is AI. The human tester supplied substantial background material, audience and purpose requirements, discussion points, an opening hook, final-thoughts structure, and a closing call to action. The tester did not rewrite the final wording. Private story details and topic are withheld.

AI04 — AI-written informational article

1,555 words. A complete AI-written informational article using headings, explanatory sections, and practical reader guidance. It tests a common search-oriented publishing format.

How we define authorship

We classify a sample according to who generated the submitted wording. Human ideas, source material, prompting, and editorial direction do not make AI-generated prose human-written. Conversely, a human document is not reclassified because its style appears formulaic. Ground truth is established before detector results are viewed, and a detector's score never changes that label.

Scroll to Top