Skip to main content
Live
Main content

New York Times says OpenAI hid evidence in ChatGPT copyright case

Court filings allege OpenAI already searched its own training data and kept a 78 million-conversation database before claiming it couldn't.

Jaeden Schafer
Editor in Chief · · 5 min read
OpenAI logo

The New York Times and The Daily News told a federal court that OpenAI has been hiding the fact that it can search its own training data and ChatGPT chat logs for their copyrighted journalism, escalating a two-year copyright lawsuit into a discovery-sanctions fight. The plaintiffs allege OpenAI already built and used a database of about 78 million de-identified ChatGPT conversations to study infringement internally, then told the court such searches would be technically burdensome and privacy-invasive. They want the judge to strip OpenAI of the ability to argue its own chat-log sample is representative, and to make the company pay their legal fees.

The claim rests on an April court-ordered deposition of OpenAI data privacy engineer Vinnie Monaco. According to the filing, Monaco testified that OpenAI had already conducted internal searches and evaluations of its training corpus for copyrighted works, and had assembled the 78 million-conversation database before the Times filed suit. Shortly after the lawsuit landed, the plaintiffs say, OpenAI implemented a Bloom filter as part of an internal toolset called Project Giraffe that detected and logged regurgitation in model outputs.

Those two revelations cut against OpenAI's core discovery posture. Throughout the case, the company has argued it cannot practically search its training data and that producing ChatGPT conversation logs would require costly retrieval, processing, and de-identification. If the plaintiffs' account of the deposition holds up, OpenAI had already done a version of that work — and kept the results.

Key facts

  • 01The Times alleges OpenAI amassed a database of 78 million de-identified ChatGPT conversations to internally evaluate copyright infringement.
  • 02Plaintiffs originally sought 120 million chat logs; OpenAI negotiated the sample down to 20 million, which the court called 'unusable' due to redactions.
  • 03OpenAI allegedly deleted billions of ChatGPT outputs after the suit was filed, in violation of a court preservation order.
  • 04An April deposition of OpenAI privacy engineer Vinnie Monaco allegedly revealed a Bloom filter tool inside 'Project Giraffe' that logged regurgitation.
  • 05The Times is asking the court to bar OpenAI from using its 20 million-log sample as evidence and to make it pay plaintiffs' legal fees.

The sample dispute compounds the problem. The plaintiffs originally asked for 120 million chat logs; OpenAI negotiated the request down to 20 million and submitted that sample last December. The court described the delivered sample as "unusable" because of the volume of redactions. The Times and Daily News further claim OpenAI deleted billions of ChatGPT outputs after the suit was filed, in violation of a court preservation order, and swapped in substitute logs for millions of records in the requested sample.

The outlets are now asking for sanctions rather than just more discovery. Their motion asks the judge to bar OpenAI from introducing the 20 million-log sample as evidence, to treat it as established fact that ChatGPT logs would have shown substantial regurgitation and grounding of the plaintiffs' content, to prohibit OpenAI from arguing the produced logs disprove regurgitation, and to shift fees.

If OpenAI genuinely believed that copying our clients' journalism was fair and legal, it wouldn't have hid the truth about having done it.
Ian B. Crosby, Lead counsel for the plaintiffs

OpenAI disputes the account. Spokesperson Drew Pusateri accused the Times of pushing to expose private user conversations as its underlying copyright claims narrow, and said the company will keep defending fair use. The Times has dropped some claims against OpenAI during the litigation, though the core allegation — that OpenAI trained on and reproduced Times journalism without a license — remains live.

The stakes of the discovery fight are larger than one publisher's case. OpenAI's public defense against every training-data copyright suit has leaned on two arguments: fair use on the merits, and the practical impossibility of auditing what a frontier model has absorbed. Project Giraffe, if it does what the plaintiffs say it does, undercuts the second argument for this specific plaintiff and this specific corpus, which is exactly what discovery is supposed to test.

As the Times' case weakens and they've been forced to drop claims against us, they're persisting with their efforts to invade the privacy of people who have nothing to do with this case, including by making these blatantly false allegations.
Drew Pusateri, OpenAI spokesperson

It also puts pressure on the other pending suits. News publishers, book authors, music labels, and image libraries are all in various stages of litigation against OpenAI and its peers, and every one of them has been told at some point that searching the training data is infeasible. A finding that OpenAI ran internal searches while telling the court it couldn't would be cited in every one of those dockets within days.

Related · from this week
OpenAI faces sanctions bid after allegedly hiding 78M ChatGPT logs from NYT
Jaeden Schafer · 5 min read →

There is a counterweight worth stating plainly. An allegation in a sanctions motion is not a court finding. OpenAI has not yet filed its formal response to this motion, and the deposition testimony the plaintiffs cite has not been publicly released in full. Judges routinely narrow sanctions requests, and a Bloom filter that flags regurgitation in outputs is not the same object as a fully searchable index of the training set — the two systems answer different questions and OpenAI is likely to argue the distinction matters.

For OpenAI, the near-term risk is not a copyright verdict — that's still far off — but an evidentiary ruling that treats the plaintiffs' worst-case theory of regurgitation as established fact for the rest of the trial. That kind of adverse-inference sanction would functionally decide the liability question before a jury ever sees it, and would set a template every other plaintiff suing a frontier lab over training data will try to copy. The technical question of whether models memorize copyrighted text is being litigated one deposition at a time, and this one landed hard.

ShareXLinkedInEmail
AI Box

Every AI model. One chat.

The latest models from ChatGPT, Claude, Gemini, Sora, ElevenLabs — 80+ models in a single chat. Compare answers side by side. Pick the best one every time.

  • ChatGPT, Claude, Gemini, Grok, DeepSeek — in one chat
  • Generate images & video with Sora, Veo, Ideogram
  • Compare any two models side by side
  • From $8.99/mo · 80+ models, all included
Try AI Boxaibox.ai
Trusted by 3,000+ teams
Got a tip?

Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.

Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.

AI Box Daily briefingFree · Daily · No fluff

Stay ahead of everyone in AI.

The tightly edited AI news email engineers, founders, and investors actually open. One email. Every weekday. Five minutes to finish.

Loved by 10,000+ AI professionals
Free forever. Unsubscribe with one click.

The briefing read inside teams at

Keep reading

More from Security

OpenAI logo
Security

OpenAI faces sanctions bid after allegedly hiding 78M ChatGPT logs from NYT

News plaintiffs say OpenAI concealed pre-searched log samples for two years while claiming the searches were technically infeasible.

Jaeden Schafer5 min read
OpenAI logo
Security

Trump administration files brief backing OpenAI in New York Times copyright suit

A 20-page DOJ filing argues fair use covers LLM training, siding with OpenAI against The New York Times in the Southern District of New York.

Jaeden Schafer4 min read
Microsoft logo
Security

NYT amends OpenAI suit, targets Microsoft's bespoke training supercomputer

The Times reframes its contributory infringement claim after a Supreme Court ruling for Cox, alleging Microsoft built the system to train on its articles.

Jaeden Schafer5 min read