A senior Microsoft executive privately described AI training practices as 'an astonishing theft of unprecedented proportions' and 'the largest theft of labor in human history,' according to unredacted filings in the three-year-old copyright lawsuit The New York Times brought against OpenAI and Microsoft. The same filings show Microsoft's own data measuring a 93% drop in click-through rates to nytimes.com from Copilot compared with traditional Bing search, and OpenAI's own leadership acknowledging that chatbots pose an 'existential threat' to publishers.
The numbers describing the scale of copying are what turn the case from a legal abstraction into a discovery problem for OpenAI and Microsoft. OpenAI's mid-training datasets alone contain more than 91,692 copies of works published by the NYT, Daily News, and Center for Investigative Reporting. A Common Crawl-derived dataset included more than 2 million documents from nytimes.com. A joint Microsoft-OpenAI initiative called Project Mango assembled a training corpus with at least 160,903 unique works from news publishers.
The internal Microsoft assessments quoted in the filing are unusually blunt. A January 2024 presentation by Brent Hecht, Microsoft's Director of Applied Science, described the traffic collapse as a 'doom loop' that would 'hurt the performance of our models and the entire web at the same time.' A separate Microsoft document warned there is a 'real risk' that generative AI could 'significantly disrupt the employment of the very people who generated the data on which the foundation model was trained.'
“It is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its 'content supply chain.'”— Brent Hecht, Microsoft Director of Applied Science
Key facts
- 01Microsoft's Copilot cut click-throughs to nytimes.com by as much as 93% compared with traditional Bing search, per internal Microsoft data.
- 02OpenAI's mid-training datasets contained more than 91,692 copies of works from the NYT, Daily News, and Center for Investigative Reporting.
- 03A Common Crawl-derived dataset used by OpenAI included more than 2 million documents from nytimes.com alone.
- 04Microsoft Director of Applied Science Brent Hecht called AI scraping 'the largest theft of labor in human history' in a January 2023 internal memo.
- 05Project Mango, a joint OpenAI-Microsoft training initiative, assembled at least 160,903 unique works from news publishers.
The language matters because it cuts against the fair-use defense OpenAI and Microsoft are running. Fair use turns in part on whether the new work substitutes for the market of the original. Microsoft CEO Satya Nadella testified in a deposition earlier this year that talking to a chatbot 'has substituted … giving you the information right there on the website on the AI platform versus needing to go to the underlying source.' OpenAI's Head of ChatGPT, Nick Turley, wrote in internal communication that products like the chatbot are 'largely substitutive' and 'will get more and more substitutive as they get better.' OpenAI President Greg Brockman described the models as 'excellent at news.'
Nadella went further in the deposition, telling lawyers that if he 'had been made aware that OpenAI had scraped and trained on information that was behind a paywall,' he would have 'invoked [Microsoft's right to] require OpenAI to retrain its models.' That statement plants the CEO of OpenAI's largest backer on the record saying paywalled content should be licensed, not scraped — a position materially different from the one OpenAI's lawyers are defending in court.
“anything that is paywalled should be licensed by anyone who wants to use it…for grounding or training”— Satya Nadella, Microsoft CEO
The filing also details how the training corpora were assembled. 'OpenAI delivered the entire GPT-3 training dataset to Microsoft, which Microsoft used to evaluate how to implement OpenAI's models within its own commercial products,' the brief reads, adding that Microsoft supplied training data to OpenAI through initiatives called Project Taxi and Project Mango. OpenAI researchers also allegedly built training datasets called WebText and WebText2 that leaned heavily on scraped news content, and pulled millions of articles from the open Common Crawl repository.
One exchange in the filing captures the tone. When OpenAI researcher Nick Ryder told Brockman about a 'hack to get around nytimes paywall,' Brockman replied 'ah nice.' The filing also alleges that copyright notices were deliberately stripped from training data before it reached the model, because researchers 'wouldn't want model outputting' copyright notices to users.
Context matters here. Much of the new material comes from The Times' own brief rather than the underlying exhibits, which remain sealed, so the quotations are presented without full context. Judges have so far been largely receptive to AI companies' fair-use arguments, and earlier this month the Trump administration filed a brief backing OpenAI's position that unlicensed training on copyrighted material should qualify as fair use. OpenAI and Microsoft did not respond to requests for comment. The disclosures are allegations and executive statements captured in discovery, not findings of liability.
Still, the substantive shift is that the plaintiffs no longer have to argue market harm from theory. They have Microsoft's own click-through telemetry showing a 93% collapse, Microsoft's CEO on the record calling for licensing of paywalled content, and OpenAI's own product leadership describing chatbots as substitutive for the underlying journalism. Those are precisely the factual elements a fair-use defense is weakest against.
For the AI industry, the filing narrows the range of viable outcomes. A settlement that includes retroactive licensing payments and a durable licensing framework for news content is now the most rational path for both defendants, particularly given Nadella's own deposition testimony. A trial verdict against OpenAI on the substitution prong would reshape how every frontier lab sources training data, and would make the licensing deals Anthropic, Google, and OpenAI itself have already signed with select publishers look like the floor rather than the ceiling. Either way, the era in which 'we scraped it and it's fair use' functioned as a complete answer is closing.
Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.
Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.



