The first AID benchmark contains four verified human-written documents and four verified AI-written documents. We use the complete master files, not detector-selected excerpts. Descriptions are intentionally broad enough to protect private source material while still showing differences in genre, length, and production method.
H01 — Human long-form narrative
3,017 words. An original human-written long-form manuscript sample. It tests whether a detector can recognize sustained personal voice and narrative development without falsely labeling it as AI.
H02 — Human nonfiction book material
3,937 words. Original human-written material about mindset and professional experience. It includes personal perspective, structured explanation, and long-form instructional prose.
H03 — Human doctoral writing
2,557 words. Original human-written doctorate-level academic material. It represents formal research-oriented prose, where standardized language can sometimes resemble machine-generated text.
H04 — Human video-script writing
1,442 words. An original human-written YouTube script. It adds conversational pacing, audience address, and spoken-language structure to the human baseline.
AI01 — AI-written professional article
2,603 words. A complete AI-written article created for a professional audience. It tests polished explanatory prose in a conventional article format.
AI02 — AI-written long-form book material
2,557 words. Complete AI-written long-form material about AI and search. It represents structured, information-dense book-style writing.
AI03 — AI-written script with extensive human direction
1,789 words, 2025. Claude generated all submitted prose, so the ground truth is AI. The human tester supplied substantial background material, audience and purpose requirements, discussion points, an opening hook, final-thoughts structure, and a closing call to action. The tester did not rewrite the final wording. Private story details and topic are withheld.
AI04 — AI-written informational article
1,555 words. A complete AI-written informational article using headings, explanatory sections, and practical reader guidance. It tests a common search-oriented publishing format.
We classify a sample according to who generated the submitted wording. Human ideas, source material, prompting, and editorial direction do not make AI-generated prose human-written. Conversely, a human document is not reclassified because its style appears formulaic. Ground truth is established before detector results are viewed, and a detector's score never changes that label.