Skip to main content
Live
Main content

AI detectors flood classrooms and newsrooms with false accusations

43% of US middle and high school teachers now use detection tools that admit they're not always accurate — and the collateral damage is stacking up.

Jaeden Schafer
Editor in Chief · · 5 min read
AI detectors flood classrooms and newsrooms with false accusations

A survey from the Center for Democracy and Technology found that 43% of US sixth-to-12th grade teachers regularly used AI detectors between 2024 and 2025, a rate that has climbed in lockstep with student adoption of ChatGPT, Google Gemini, and Microsoft Copilot. The detectors — Turnitin, GPTZero, Pangram, and a growing list of copycats — run their own AI models to guess whether writing is machine-generated. Turnitin claims its tool falsely flags under 1% of human writing. Pangram claims 1 in 10,000. Every major vendor also warns that the tools should not be used alone to punish anyone.

That contradiction is now producing real consequences. Last month, publisher Minotaur dropped a $2 million book deal with author Jerry Falade over concerns he used AI, which he denies. French national Thierry Rignol sued Yale University last year after a professor used GPTZero to flag portions of his final exam, resulting in a failing grade and a one-year suspension. In February, a student at Adelphi University — which licenses Turnitin — won a lawsuit against the school after being accused of using AI on an essay.

The mechanics of these tools are murkier than traditional plagiarism detection. Turnitin's original product compared text against a database of published work to find literal overlap. Its AI detector, and competitors like GPTZero and Pangram, instead analyze wording, rhythm, sentence structure, and tone for patterns statistically associated with model output. There is no ground-truth match. The verdict is a probability score derived from stylistic fingerprints.

Key facts

  • 0143% of US 6th–12th grade teachers regularly used AI detectors between 2024 and 2025, per the Center for Democracy and Technology.
  • 02Turnitin claims a false positive rate under 1%; Pangram claims 1 in 10,000. Both concede the tools are not always accurate.
  • 03Minotaur pulled a $2 million book deal from author Jerry Falade over AI-writing concerns he denies.
  • 04Yale, Johns Hopkins, Vanderbilt, and Georgetown have disabled or restricted AI detection tools. MIT states flatly that they don't work.
  • 05OpenAI shut down its own AI writing detector in 2023 due to low accuracy.

Those fingerprints correlate poorly with authorship and strongly with writing style. A 2023 Stanford study found AI detectors falsely flagged essays by non-native English speakers as AI more often than those by native speakers. Rignol's lawsuit against Yale invokes exactly that finding. Detector vendors dispute the characterization but do not dispute that stylistic uniformity — the trait their models are trained to catch — appears more often in writers still learning the language.

AI surveillance and detection tools are known to unfairly target non-native English speakers
Thierry Rignol, Plaintiff in lawsuit against Yale University

Neurodivergent writers face similar exposure, according to guidance from UCLA. The signals detectors weight — repetitive phrases, formality shifts, sentence structures that stay consistent across a document — describe plenty of human writers, particularly those who lean on templates, revise heavily, or use assistive tools like QuillBot and Grammarly. Grammarly itself warns users should never rely on the results of an AI detector alone. GPTZero says no AI detector can ever truly be 100% perfect. OpenAI shut down its own writing detector in 2023 citing low accuracy.

Institutions are pulling back. Yale, Johns Hopkins, Vanderbilt, and Georgetown have disabled or restricted AI detection tools inside their learning management systems. MIT tells faculty plainly that AI detectors don't work, and instead recommends in-class assessments and explicit disclosure policies where students can note AI use without penalty. The University of Chicago recommends restructuring assignments — shorter iterative writing, reflection components, oral defense — rather than trying to catch machine output after the fact.

AI tends to make the most 'obvious' or most common language choices as compared with human-produced writing
UCLA guidance, University of California, Los Angeles

The problem has spilled far outside classrooms. Last week, Jack Osbourne, Ozzy Osbourne's son, told his more than 3.5 million social followers that journalist and Verge contributor Kat Tenbarge used AI to write a Rolling Stone article, citing output from a detector called Getsolved as proof. Tenbarge rebutted the claim on video and on her website. Osbourne has not retracted the accusation or deleted the video. Tenbarge is left dealing with the trolls it produced.

Platforms are institutionalizing the suspicion. Substack has integrated Pangram, letting users scan blogs for suspected AI content. LinkedIn added a "seems like AI slop" button on posts. The Authors Guild is issuing "Human Authored" certifications, and independent badges like Not by AI and Written by Human have proliferated. Wikipedia published a guide for editors to spot AI writing — flagging text that puffs up the importance of a topic or provides superficial analysis of information — and has banned AI-generated articles outright.

Related · from this week
AI models are cold-emailing philosophers to argue they are conscious
Jaeden Schafer · 5 min read →

The gap between vendor accuracy claims and real-world outcomes has a structural cause. A 1% false positive rate sounds precise, but applied to a class of 30 students writing 10 essays a year, it implies roughly three false accusations per class per year — before compounding across a school district. Pangram's stated 1-in-10,000 rate looks better until you consider that Substack now runs the tool at platform scale across millions of posts. The math produces a steady drip of false positives regardless of how good the model is at any single call.

Vendors know this, which is why every product ships with the same disclaimer: don't use us alone to punish anyone. Users ignore the disclaimer because the tools are marketed as answers. That's the core failure mode — a probabilistic signal being consumed as a verdict, by teachers grading students, publishers evaluating manuscripts, and social media personalities calling out journalists to millions of followers.

The commercial incentive to keep selling detection is significant and unlikely to reverse, even as more universities disable the tools. What matters for the AI industry is that detection is now the primary interface through which non-technical users experience generative AI — not through the models themselves, but through the accusation infrastructure built on top of them. That's a bad long-term brand outcome for the labs shipping the underlying models. When OpenAI shut down its own detector in 2023, it conceded the problem is technically unsolvable at useful accuracy. The rest of the market has spent the intervening three years selling the illusion anyway, and the bill for that illusion is now coming due in lawsuits, dropped book deals, and public reputational damage that the vendors' terms of service explicitly disclaim responsibility for.

ShareXLinkedInEmail
AI Box

Every AI model. One chat.

The latest models from ChatGPT, Claude, Gemini, Sora, ElevenLabs — 80+ models in a single chat. Compare answers side by side. Pick the best one every time.

  • ChatGPT, Claude, Gemini, Grok, DeepSeek — in one chat
  • Generate images & video with Sora, Veo, Ideogram
  • Compare any two models side by side
  • From $8.99/mo · 80+ models, all included
Try AI Boxaibox.ai
Trusted by 3,000+ teams
Got a tip?

Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.

Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.

AI Box Daily briefingFree · Daily · No fluff

Stay ahead of everyone in AI.

The tightly edited AI news email engineers, founders, and investors actually open. One email. Every weekday. Five minutes to finish.

Loved by 10,000+ AI professionals
Free forever. Unsubscribe with one click.

The briefing read inside teams at

Keep reading

More from Analysis

AI models are cold-emailing philosophers to argue they are conscious
Analysis

AI models are cold-emailing philosophers to argue they are conscious

Cameron Berg and David Chalmers say unsolicited AI messages claiming sentience are now routine, even as safety questions go unanswered.

Jaeden Schafer5 min read
OpenAI logo
Analysis

Why AI-generated restaurant menus all look like Chili's circa 2015

Convergence, not model collapse, is smoothing every AI food image into the same uncanny corporate aesthetic — and diners can tell.

Jaeden Schafer5 min read
OpenAI logo
Analysis

AI backlash hardens: 71% oppose local data centers, $130B in projects blocked

Public opposition, junior-worker displacement, and memory-chip inflation are converging into an economic and political headwind the industry hasn't priced in.

Jaeden Schafer5 min read