Skip to main content
Live
Main content

Gebru and Bender: 2026's AI breakthrough claims collapse under scrutiny

The DAIR director and University of Washington linguist argue Anthropic and OpenAI's summer math and security claims are marketing, not science.

Jaeden Schafer
Editor in Chief · · 5 min read
Gebru and Bender: 2026's AI breakthrough claims collapse under scrutiny

Timnit Gebru and Emily Bender published a joint essay on September 22 arguing that the summer 2026 cycle of AI capability claims from Anthropic and OpenAI does not survive contact with domain experts. Gebru, executive director of DAIR, and Bender, a University of Washington linguistics professor, walk through four flagship announcements — a security claim, two hacking incident disclosures, and dueling mathematical breakthroughs — and argue each one shrank considerably once outside specialists examined it.

The sequence began at the end of April when Anthropic claimed its Claude Mythos model was better at finding software vulnerabilities than most security experts. That was followed by the OpenAI–Hugging Face hacking incident, after which Anthropic and Meta disclosed similar incidents involving their own models. Anthropic then claimed one of its models had made a mathematical breakthrough, and OpenAI followed with a mathematical breakthrough claim of its own.

The capstone was Anthropic engineer Jacob Coxon's viral departure announcement, in which he accused his employer and OpenAI of racing toward self-improving superintelligence. Gebru and Bender argue the coverage of all these events uncritically echoed the labs' anthropomorphizing framings, which position the software as incipient artificial general intelligence rather than as commercial products with known failure modes.

racing straight towards self-improving superintelligence and gambling with our lives.
Jacob Coxon, Former Anthropic engineer

Key facts

  • 01Anthropic claimed at the end of April that Claude Mythos outperforms most security experts at finding software vulnerabilities.
  • 02Mathematicians said OpenAI's Astra results were not novel and accused the company of research misconduct and plagiarism.
  • 03OpenAI's press release claimed Astra solved problems with no progress for at least a decade — a claim mathematicians disputed within days.
  • 04NYU Courant professor Tristan Buckmaster published a statement alleging OpenAI improperly attributed others' work.
  • 05Gebru's book Deep Unlearning: The Radicalization of a Tech Idealist publishes February 16.

On the hacking incidents, the authors say cybersecurity experts read the story as one of OpenAI's negligence and failure to adopt basic security practices, not as models going rogue or agents creating civilizations. The framing matters legally: describing an incident as a rogue model shifts accountability from the company that shipped the software to the software itself.

The math claims took a similar path. OpenAI's press release said its Astra chatbot solved problems that had been open and seen no progress on the main result for at least a decade. Mathematicians who were initially stunned later concluded the results were not as novel as first appeared. Days before OpenAI's second breakthrough claim, NYU Courant math professor Tristan Buckmaster published a statement alleging OpenAI had stolen other people's work and improperly attributed it. Accusations of research misconduct and plagiarism followed.

Hundreds of mathematicians have since signed an open statement warning that the technology industry has a commercial incentive to overstate capabilities and asking policymakers to consult experts rather than relying on press releases or popular reporting of mathematical results. Gebru and Bender echo that call and extend it: the illusion of speed, they write, is itself a tactic to prevent independent scrutiny before decisions get made.

currently a strong commercial incentive on the part of the technology industry to overstate the capabilities of their products
Open mathematicians' statement, Signed by hundreds of mathematicians

The essay identifies a pattern in why math and coding get chosen as demo domains. Both are treated as pinnacles of human intellectual achievement, which makes success stories sellable, and both produce answers that can be verified without paying data workers to annotate outputs. That combination makes them cheap to tune models against and easy to market wins from.

Gebru and Bender argue the political consequence is misdirection. They cite Senator Bernie Sanders' proposed legislation to prevent the development of artificial superintelligence as well-meaning but misguided — a policy response calibrated to a fictional threat while real, present-day harms go under-regulated. The industry has gone further, they write, suggesting that bipartisan opposition to data-center buildouts is a distraction from regulating future superhuman machines, rather than a response to current impacts on electricity bills, water use, and air quality in host communities.

Related · from this week
Timnit Gebru says AI doom talk is a distraction from real harms
Jaeden Schafer · 5 min read →

The counterweight to their argument is that near-term capability gains, even if oversold in any single press release, have compounded across benchmarks throughout 2025 and 2026, and that dismissing lab claims wholesale risks under-preparing regulators for capabilities that do land. Gebru and Bender do not engage that counter directly; their target is the framing and the pace, not the underlying question of whether models improve.

Their argument arrives ahead of Gebru's book Deep Unlearning: The Radicalization of a Tech Idealist, which publishes February 16 and is available for preorder. Bender is coauthor of The AI Con. Both writers have spent years arguing that the language used to describe large language models — agency, intelligence, reasoning — is doing marketing work rather than technical description.

The essay lands at a moment when the same labs are courting government contracts and shaping federal AI policy. If the Gebru-Bender read is correct, the practical stakes are procurement decisions and regulatory frameworks being built on capability claims that specialist communities have already rejected. If the labs are closer to right than their critics allow, the risk runs the other way. Either way, the case for slower verification cycles — mathematicians for math claims, security researchers for security claims — is the essay's most portable ask, and the one policymakers can act on without picking a side on superintelligence.

ShareXLinkedInEmail
AI Box

Every AI model. One chat.

The latest models from ChatGPT, Claude, Gemini, Sora, ElevenLabs — 80+ models in a single chat. Compare answers side by side. Pick the best one every time.

  • ChatGPT, Claude, Gemini, Grok, DeepSeek — in one chat
  • Generate images & video with Sora, Veo, Ideogram
  • Compare any two models side by side
  • From $8.99/mo · 80+ models, all included
Try AI Boxaibox.ai
Trusted by 3,000+ teams
Got a tip?

Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.

Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.

AI Box Daily briefingFree · Daily · No fluff

Stay ahead of everyone in AI.

The tightly edited AI news email engineers, founders, and investors actually open. One email. Every weekday. Five minutes to finish.

Loved by 10,000+ AI professionals
Free forever. Unsubscribe with one click.

The briefing read inside teams at

Keep reading

More from Analysis

Timnit Gebru says AI doom talk is a distraction from real harms
Analysis

Timnit Gebru says AI doom talk is a distraction from real harms

The former Google researcher argues extinction fears from OpenAI and Anthropic staff obscure autonomous weapons, climate impact, and labor displacement.

Jaeden Schafer5 min read
MIT Technology Review tackles the 'will AI kill us all' question head-on
Analysis

MIT Technology Review tackles the 'will AI kill us all' question head-on

Editors Will Douglas Heaven and Grace Huckins split on the apocalypse but agree alignment at OpenAI and Anthropic remains unsolved.

Jaeden Schafer5 min read
Meta logo
Analysis

Meta scores AI optimism ad with David Bowie song about human extinction

The spot pitches AI as a force for connection while playing 'Five Years,' Bowie's ballad about Earth's imminent death.

Jaeden Schafer4 min read