Skip to main content
Live
Main content

Mathematicians accuse OpenAI of using unpublished work in math breakthroughs

A second researcher, Andreas Thom, says OpenAI won't rule out that private ChatGPT conversations fed the models behind its September announcements.

Jaeden Schafer
Editor in Chief · · 5 min read
OpenAI logo

A second mathematician has publicly accused OpenAI of dishonesty over the training data behind its recent string of math breakthroughs, deepening a dispute that began days earlier over whether the company's models benefited from unpublished user work. Andreas Thom, whose specialty is non-sofic groups, said in a series of Mastodon posts on September 10, 2026 that OpenAI has refused to conclusively rule out that his own interactions with ChatGPT contributed to one of the 10 results the company announced last month with what he described as great fanfare.

Non-sofic groups are, roughly, infinite mathematical structures that cannot be approximated by finite ones. OpenAI's announced result in that area drew immediate criticism in mathematical circles for failing to credit recent work by Thom and Gábor Kun, and the company quietly amended its writeup after the pushback. Thom said what struck him most was OpenAI's detailed command of our techniques — approaches that were neither the most obvious nor the most promising routes to a solution at the time.

Thom said he emailed OpenAI researchers Sébastien Bubeck and Mark Sellke, the latter also a statistician at Harvard, to ask whether his conversations with ChatGPT were part of the training data or accessible to the reasoning process. The response, he said, addressed only whether the chatbot could directly retrieve his conversations, not whether they had entered the broader pools of training data OpenAI uses to improve its models.

Key facts

  • 01Mathematician Andreas Thom accused OpenAI of 'dishonesty' over the origins of the training data behind its 10 math results announced last month.
  • 02Thom's specialty, non-sofic groups, was one of the 10 results; OpenAI quietly amended its writeup after failing to credit Thom and Gábor Kun.
  • 03OpenAI told Thom via researchers Sébastien Bubeck and Mark Sellke only that his ChatGPT conversations weren't directly accessed, not whether they entered training data.
  • 04The complaint follows NYU professor Tristan Buckmaster questioning whether his use of Codex fed OpenAI's Millennium Prize Navier-Stokes result.
  • 05OpenAI said it 'cannot rule out that de-identified data' from user products helped improve its models.

That gap is the crux of the complaint. Thom argues researchers have no way to reverse-engineer OpenAI's training pipeline: Only OpenAI has the relevant data for that. If the company wants to deny using user research, he said, the burden falls on it to disclose the datasets and settings that govern how customer inputs are handled.

The dispute follows a parallel row involving Tristan Buckmaster, a mathematics professor at New York University who publicly questioned whether OpenAI's models had benefited from his use of Codex. Buckmaster had been working on the same problems that OpenAI later announced it had solved — including a Millennium Prize breakthrough on the Navier-Stokes equations, which describe the movement of fluids — alongside Anthropic researcher Levent Alpöge in a personal capacity.

In its Navier-Stokes announcement, OpenAI said the researchers and the agents did not see any of their work through any means until they released it publicly, and that no specific user data was accessed in order to solve this problem. The company then added a hedge: while unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models. Thom said that same distinction was drawn in his own correspondence with Bubeck and Sellke.

De-identification may remove a name; it does not remove the intellectual content of a mathematical idea.
Andreas Thom, Mathematician

Thom's judgment on the exchange was sharp: Sellke's categorical answer was, at minimum, unjustifiably broad and materially misleading; looking back it was plainly dishonest. He added that it would be ethically indefensible if nonpublic research supplied by users helped train models that OpenAI then used to race those same users to publication without consent, disclosure, or credit.

OpenAI did not immediately respond to a request for comment on Thom's statements. The company has not published a technical accounting of which datasets, user conversations, or product telemetry flowed into the models it deployed for the math results, and its terms of service around business and consumer ChatGPT tiers differ in how user inputs may or may not be retained for training.

Related · from this week
OpenAI moves to dismiss Apple trade secrets suit, calls it 'rotten to its core'
Jaeden Schafer · 4 min read →

The disputes are unusual because they arrive during what should be a moment of validation for OpenAI's research program. The Navier-Stokes result, if verified, is a genuine mathematical achievement — one of a small number of Millennium Prize problems, each carrying a $1M reward for a correct solution. But the surrounding conduct has soured the reception in the field, and researchers have told reporters they worry mathematicians will now work more secretly, wary that even rumors of progress could trigger a race with a well-funded lab.

The training-data question is the same one that has driven copyright suits from authors, publishers, and the New York Times against OpenAI: what went into the model, when, and with whose permission. Mathematical research adds a new wrinkle because the inputs in question are not published books scraped from the open web but private, in-progress ideas shared inside a paid product by researchers who assumed the conversation was confidential. The distinction between direct retrieval and indirect training influence — the one OpenAI keeps drawing — matters enormously to lawyers and almost not at all to a researcher whose unpublished technique appears in someone else's proof.

For OpenAI, the reputational cost of not disclosing more is beginning to compound. Each new breakthrough now arrives with a question attached about where the underlying ideas came from, and every non-denial hedge — 'we cannot rule out' — invites the next mathematician to check their chat history. The company can defuse this cheaply by publishing what data trained the reasoning models used for the math work, or expensively by continuing to answer in a register that its own users are now describing as dishonest. So far, it has chosen the second.

ShareXLinkedInEmail
AI Box

Every AI model. One chat.

The latest models from ChatGPT, Claude, Gemini, Sora, ElevenLabs — 80+ models in a single chat. Compare answers side by side. Pick the best one every time.

  • ChatGPT, Claude, Gemini, Grok, DeepSeek — in one chat
  • Generate images & video with Sora, Veo, Ideogram
  • Compare any two models side by side
  • From $8.99/mo · 80+ models, all included
Try AI Boxaibox.ai
Trusted by 3,000+ teams
Got a tip?

Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.

Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.

AI Box Daily briefingFree · Daily · No fluff

Stay ahead of everyone in AI.

The tightly edited AI news email engineers, founders, and investors actually open. One email. Every weekday. Five minutes to finish.

Loved by 10,000+ AI professionals
Free forever. Unsubscribe with one click.

The briefing read inside teams at

Keep reading

More from Security

OpenAI logo
Security

OpenAI moves to dismiss Apple trade secrets suit, calls it 'rotten to its core'

OpenAI's motion says Apple mischaracterized 'generic' product work as trade secrets; the judge hears arguments October 1st.

Jaeden Schafer4 min read
OpenAI logo
Security

OpenAI subpoenaed by New York AG in multistate investigation

The probe spans advertising, model sycophancy, consumer data, and treatment of minors — and lands days after OpenAI filed confidentially to IPO.

Jaeden Schafer5 min read
OpenAI logo
Security

Florida sues OpenAI and Sam Altman over ChatGPT-linked violence

The first state-led suit against OpenAI ties ChatGPT to multiple murders and seeks damages under Florida's unfair trade laws.

Jaeden Schafer5 min read