Microsoft told a federal court on September 4, 2026 that its Copilot assistant almost never reproduces meaningful chunks of New York Times articles or copyrighted books, citing an analysis of 8.2 million chat logs it handed over in discovery. Fewer than 1% of those logs — 59,545 conversations — contained even 16 words in common with news content used to ground the model. Microsoft is using the numbers to push for summary judgment, an early exit from the consolidated copyright lawsuit brought by The New York Times, the Center for Investigative Reporting, and the Authors Guild.
The 8.2 million logs were not a random sample. Microsoft says they were selected precisely because they hit on keywords implicating the plaintiffs' websites, meaning they were the conversations most likely to surface copyrighted material. Even inside that stacked dataset, an expert for the Center for Investigative Reporting found 51 instances of substantial overlap with CIR work — a rounding error against millions of chats.
The results for book authors are thinner still. Across the same 8.2 million conversations, an expert in the authors' suit found only 24 responses that contained at least 30 matching words from the plaintiffs' books. Of 212 books evaluated, just 10 produced any matches at all. Microsoft argues those figures rebut the core theory that Copilot functions as a substitute product for the underlying works.
Key facts
- 01Microsoft handed over 8.2 million Copilot chat logs in discovery, selected because they were most likely to contain plaintiffs' works.
- 02Of those logs, 59,545 contained at least 16 words in common with news content used to ground the model — under 1% of the dataset.
- 03Only 24 responses across 8.2 million conversations contained at least 30 matching words from author plaintiffs' books.
- 04Just 10 of 212 books evaluated had any matches at all, according to Microsoft's filing.
- 05Microsoft filed the analysis on September 4, 2026 in support of a summary judgment motion; the Trump administration filed a statement of interest supporting OpenAI in the same case.
Microsoft's legal frame is fair use. The company argues that training large language models on copyrighted text is transformative because the resulting systems are used for purposes wholly unlike the originals — writing code, drafting emails, answering questions. Occasional reproduction of source text, Microsoft told the court, "hardly undermines the transformative purpose of LLM training." That is the fair-use standard the company is asking the judge to apply at summary judgment.
The Times rejects the framing outright. Lead counsel Ian Crosby said discovery has established that Microsoft and OpenAI built commercial products designed to substitute for Times journalism and compete directly with its business. The plaintiffs have consistently argued that the volume of reproduction understates the harm: even partial substitution, at scale, siphons readers and advertising revenue from the outlets whose reporting fed the training data.
The consolidated case bundles the news publishers' and book authors' claims under a single judge — a procedural move the plaintiffs opposed, since it forces two distinct theories of harm through the same legal filter. Microsoft's Friday filing leans into that consolidation, presenting both datasets side by side to argue that the fact patterns look similar: high-volume access to copyrighted text, low-volume regurgitation at the output.
Publisher lawsuits have become the central legal risk for every frontier lab shipping general-purpose chatbots. Anthropic settled with the Authors Guild in a separate action earlier this year. OpenAI faces parallel suits from multiple newspaper chains. A summary judgment in Microsoft's favor here would not bind those other cases directly, but it would give every AI defendant a template — a discovery dataset, a matching-word analysis, and a fair-use argument tied to concrete numbers rather than abstract claims about model behavior.
There are open questions the filings do not resolve. Sixteen matching words is a low bar for identifying overlap but a high bar for proving substitution — a chatbot answer that echoes a Times phrase is not the same as one that reproduces a full article. Conversely, the plaintiffs will argue that Microsoft's keyword-filtered sample understates real-world reproduction, because the queries in production may look nothing like the discovery set. The Trump administration also filed a statement of interest this week supporting OpenAI's position, adding federal weight to the fair-use argument.
If the judge grants summary judgment, the case ends before trial and Microsoft locks in a favorable precedent on training-data fair use. If the judge denies it, discovery continues and the plaintiffs get to test their substitution theory in front of a jury — a much riskier posture for the AI defendants, given the sums involved and the emotional weight of copyright claims from named journalists and authors. Every lab shipping a retrieval-augmented chatbot is watching for which way that ruling breaks, because it will set the operating cost of grounding models on the open web for years.
Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.
Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.




