A survey from the Center for Democracy and Technology found that 43% of US sixth-to-12th grade teachers regularly used AI detectors between 2024 and 2025, a rate that has climbed in lockstep with student adoption of ChatGPT, Google Gemini, and Microsoft Copilot. The detectors — Turnitin, GPTZero, Pangram, and a growing list of copycats — run their own AI models to guess whether writing is machine-generated. Turnitin claims its tool falsely flags under 1% of human writing. Pangram claims 1 in 10,000. Every major vendor also warns that the tools should not be used alone to punish anyone.
That contradiction is now producing real consequences. Last month, publisher Minotaur dropped a $2 million book deal with author Jerry Falade over concerns he used AI, which he denies. French national Thierry Rignol sued Yale University last year after a professor used GPTZero to flag portions of his final exam, resulting in a failing grade and a one-year suspension. In February, a student at Adelphi University — which licenses Turnitin — won a lawsuit against the school after being accused of using AI on an essay.
The mechanics of these tools are murkier than traditional plagiarism detection. Turnitin's original product compared text against a database of published work to find literal overlap. Its AI detector, and competitors like GPTZero and Pangram, instead analyze wording, rhythm, sentence structure, and tone for patterns statistically associated with model output. There is no ground-truth match. The verdict is a probability score derived from stylistic fingerprints.
Key facts
- 0143% of US 6th–12th grade teachers regularly used AI detectors between 2024 and 2025, per the Center for Democracy and Technology.
- 02Turnitin claims a false positive rate under 1%; Pangram claims 1 in 10,000. Both concede the tools are not always accurate.
- 03Minotaur pulled a $2 million book deal from author Jerry Falade over AI-writing concerns he denies.
- 04Yale, Johns Hopkins, Vanderbilt, and Georgetown have disabled or restricted AI detection tools. MIT states flatly that they don't work.
- 05OpenAI shut down its own AI writing detector in 2023 due to low accuracy.
Those fingerprints correlate poorly with authorship and strongly with writing style. A 2023 Stanford study found AI detectors falsely flagged essays by non-native English speakers as AI more often than those by native speakers. Rignol's lawsuit against Yale invokes exactly that finding. Detector vendors dispute the characterization but do not dispute that stylistic uniformity — the trait their models are trained to catch — appears more often in writers still learning the language.
“AI surveillance and detection tools are known to unfairly target non-native English speakers”— Thierry Rignol, Plaintiff in lawsuit against Yale University
Neurodivergent writers face similar exposure, according to guidance from UCLA. The signals detectors weight — repetitive phrases, formality shifts, sentence structures that stay consistent across a document — describe plenty of human writers, particularly those who lean on templates, revise heavily, or use assistive tools like QuillBot and Grammarly. Grammarly itself warns users should never rely on the results of an AI detector alone. GPTZero says no AI detector can ever truly be 100% perfect. OpenAI shut down its own writing detector in 2023 citing low accuracy.
Institutions are pulling back. Yale, Johns Hopkins, Vanderbilt, and Georgetown have disabled or restricted AI detection tools inside their learning management systems. MIT tells faculty plainly that AI detectors don't work, and instead recommends in-class assessments and explicit disclosure policies where students can note AI use without penalty. The University of Chicago recommends restructuring assignments — shorter iterative writing, reflection components, oral defense — rather than trying to catch machine output after the fact.
“AI tends to make the most 'obvious' or most common language choices as compared with human-produced writing”— UCLA guidance, University of California, Los Angeles
The problem has spilled far outside classrooms. Last week, Jack Osbourne, Ozzy Osbourne's son, told his more than 3.5 million social followers that journalist and Verge contributor Kat Tenbarge used AI to write a Rolling Stone article, citing output from a detector called Getsolved as proof. Tenbarge rebutted the claim on video and on her website. Osbourne has not retracted the accusation or deleted the video. Tenbarge is left dealing with the trolls it produced.
Platforms are institutionalizing the suspicion. Substack has integrated Pangram, letting users scan blogs for suspected AI content. LinkedIn added a "seems like AI slop" button on posts. The Authors Guild is issuing "Human Authored" certifications, and independent badges like Not by AI and Written by Human have proliferated. Wikipedia published a guide for editors to spot AI writing — flagging text that puffs up the importance of a topic or provides superficial analysis of information — and has banned AI-generated articles outright.
The gap between vendor accuracy claims and real-world outcomes has a structural cause. A 1% false positive rate sounds precise, but applied to a class of 30 students writing 10 essays a year, it implies roughly three false accusations per class per year — before compounding across a school district. Pangram's stated 1-in-10,000 rate looks better until you consider that Substack now runs the tool at platform scale across millions of posts. The math produces a steady drip of false positives regardless of how good the model is at any single call.
Vendors know this, which is why every product ships with the same disclaimer: don't use us alone to punish anyone. Users ignore the disclaimer because the tools are marketed as answers. That's the core failure mode — a probabilistic signal being consumed as a verdict, by teachers grading students, publishers evaluating manuscripts, and social media personalities calling out journalists to millions of followers.
The commercial incentive to keep selling detection is significant and unlikely to reverse, even as more universities disable the tools. What matters for the AI industry is that detection is now the primary interface through which non-technical users experience generative AI — not through the models themselves, but through the accusation infrastructure built on top of them. That's a bad long-term brand outcome for the labs shipping the underlying models. When OpenAI shut down its own detector in 2023, it conceded the problem is technically unsolvable at useful accuracy. The rest of the market has spent the intervening three years selling the illusion anyway, and the bill for that illusion is now coming due in lawsuits, dropped book deals, and public reputational damage that the vendors' terms of service explicitly disclaim responsibility for.
Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.
Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.




