The New York Times is asking a federal court to sanction OpenAI for allegedly concealing evidence in the copyright case over ChatGPT's use of news content, in a July 9, 2026 motion that could reshape one of the most closely watched AI lawsuits in the country. News plaintiffs led by the Times accuse OpenAI of hiding two large de-identified log samples — one of 10 million and one of 78 million ChatGPT conversations — while telling the court for two years that searching such data was technically infeasible. The samples matter because they are the primary evidence pool for whether ChatGPT users routinely prompted the model to regurgitate paywalled articles.
The alleged concealment surfaced in an April re-deposition of OpenAI privacy engineer Vincent Monaco, whom the court compelled to testify again after his first appearance was deemed inadequate. According to the Times' filing, Monaco testified that OpenAI had already searched those samples internally as part of building a filter to block regurgitation of copyrighted content. Plaintiffs say that admission contradicts two years of OpenAI representations that log searches would take months of engineering work and violate user privacy.
The Times originally sought 120 million news-related logs. Instead, it says it spent eight months confined to a 20 million log sample inside a court-supervised 'sandbox,' a scope plaintiffs argue was narrowed based on OpenAI's 'false representations regarding its existing technical capabilities.' OpenAI then applied 19 billion AI-generated redactions to that 20 million sample — so many that the court itself called the resulting dataset 'unusable.' Some redactions were later removed, but plaintiffs say domain names and publisher fields remained blacked out, hampering their searches.
Key facts
- 01The New York Times filed a July 9, 2026 motion seeking sanctions against OpenAI over allegedly concealed ChatGPT log samples of 10 million and 78 million records.
- 02News plaintiffs say OpenAI forced them to spend eight months searching a 20 million log sample, far smaller than the 120 million logs they requested.
- 03OpenAI applied 19 billion AI-generated redactions to the 20 million sample, which the court found 'unusable.'
- 04OpenAI privacy engineer Vincent Monaco testified at an April re-deposition that OpenAI already had the capability to search the larger samples.
- 05Plaintiffs are asking the court to bar OpenAI from using the 20 million sample and to instruct the jury that OpenAI deleted billions of logs.
The alleged double standard is at the center of the sanctions request. OpenAI ran regurgitation searches on the larger samples for its own product research, plaintiffs contend, while telling the court the same work was too costly to perform for discovery. The Times argues that behavior 'withheld highly relevant evidence, prolonged discovery, inflated expenses, and burdened the Court.'
“OpenAI was willing and able to search its output logs—when it benefitted OpenAI”— The New York Times, news plaintiffs' filing
Timing compounds the dispute. Near the end of discovery, OpenAI told plaintiffs that the 78 million log sample had actually been available for inspection for 'over a year' — an assertion the Times calls implausible given OpenAI's public arguments that expanding log access would violate ChatGPT user privacy. Plaintiffs frame it as a binary: either OpenAI accidentally produced the dataset and lost track of it, or it buried the production and concealed that fact while fighting to keep logs sealed.
Ian Crosby, the Times' lead counsel, put the accusation in unusually direct terms. The filing also alleges OpenAI deleted parts of the 20 million sample and, more broadly, compressed or deleted billions of logs that should have been preserved under the court's preservation order. Monaco reportedly testified that OpenAI 'thought about complying' with the order but decided the work would be too hard.
OpenAI disputes the characterization. A company spokesperson linked the sanctions motion to what it described as an eroding case, noting the Times recently dropped some claims against OpenAI. The company frames continued log requests as a privacy intrusion on ChatGPT users unrelated to the litigation, and maintains its training practices are protected by fair use.
Times spokesperson Graham James rejected the weakening-case narrative last month, telling Ars Technica that dropping certain claims streamlined rather than diminished the suit, particularly as claims against Microsoft were expanded. 'Our core claims remain the same from the day we filed this lawsuit — that Microsoft and OpenAI stole millions of The Times's copyrighted works to compete with our products and illegally enrich themselves,' James said.
“As the Times' case weakens and they've been forced to drop claims against us, they're persisting with their efforts to invade the privacy of people who have nothing to do with this case, including by making these blatantly false allegations.”— OpenAI spokesperson, OpenAI
The remedies plaintiffs are seeking would materially damage OpenAI's defense. They want the court to bar OpenAI from using the heavily redacted 20 million log sample at trial, to find as a matter of law that the withheld logs contained 'substantial' regurgitation of Times and Daily News content, and to instruct the jury that OpenAI deleted billions of logs subject to preservation. 'Lesser sanctions would not be effective,' the filing argues. OpenAI has not yet filed a formal response to the sanctions motion beyond its public statement.
For anyone tracking the economics of the AI copyright fight, this is the pivot point. The fair-use analysis in generative AI cases turns heavily on market harm, and market harm turns on whether models actually reproduce protected content at scale. If the court adopts the plaintiffs' proposed adverse inferences, OpenAI would enter trial legally presumed to have generated substantial regurgitation — a posture that makes settling far cheaper than fighting, and that would set a template every other publisher suing an AI lab will want to copy. The broader lesson for the industry is procedural rather than doctrinal: how labs handle discovery on training data and output logs may end up doing more to shape AI copyright law than any single fair-use ruling.
Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.
Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.




