Skip to main content
Live
Main content

Macmillan, McGraw Hill, and Hachette sue Meta over Llama training data

Five publishers and author Scott Turow allege Meta pulled copyrighted books from LibGen and Anna's Archive to train Llama.

Jaeden Schafer
Editor in Chief · · 5 min read
Meta logo

Five major book publishers and author Scott Turow sued Meta on May 5, 2026, alleging the company committed "one of the most massive infringements of copyrighted materials in history" while training its Llama models. Macmillan, McGraw Hill, Elsevier, Hachette, and Cengage filed the class action together, accusing Meta of knowingly pulling books and journal articles from pirate repositories and feeding them to its model. The publishers want damages, an injunction, and a full list of copyrighted works ingested during training.

The complaint names LibGen, Anna's Archive, Sci-Hub, and Sci-Mag as the alleged sources, and adds that Meta also trained Llama on the Common Crawl dataset, which the plaintiffs say is "full of unauthorized copies of copyrighted works." The publishers argue Meta knew the provenance of the material and used it anyway. Internal Meta discussions surfaced in earlier author suits referenced "media coverage suggesting we have used a dataset we know to be pirated."

The most concrete allegation is reproduction. The publishers claim Llama "outputs verbatim and near-verbatim substitutes" of copyrighted material, citing the 9th edition of Calculus: Early Transcendentals by James Stewart, a Cengage title. According to the filing, prompting Llama with two brief sentences from the textbook causes the model to continue the section word-for-word. That is the kind of memorization courts have treated as a fact pattern distinct from abstract questions about fair use.

Key facts

  • 01Macmillan, McGraw Hill, Elsevier, Hachette, Cengage, and author Scott Turow filed the class action against Meta on May 5, 2026.
  • 02The complaint alleges Meta sourced training data from pirate sites including LibGen, Anna's Archive, Sci-Hub, and Sci-Mag.
  • 03Plaintiffs say Llama reproduces verbatim continuations of the 9th edition of James Stewart's Calculus: Early Transcendentals.
  • 04Anthropic settled a parallel authors' class action last year for $1.5 billion over allegedly pirated books.
  • 05Meta spokesperson Dave Arnold said the company will 'fight this lawsuit aggressively,' citing fair use.

Meta's response is unchanged from its earlier posture. "AI is powering transformative innovations, productivity and creativity for individuals and companies, and courts have rightly found that training AI on copyrighted material can qualify as fair use," spokesperson Dave Arnold said. "We will fight this lawsuit aggressively." Meta has not addressed the verbatim-output claims publicly.

When prompted with two sentences from the 9th edition of Stewart's Calculus: Early Transcendentals, Llama reproduces the section word-for-word, the publishers allege.
Jaeden Schafer

Meta has reason for confidence on one front. Last year, a federal judge ruled in the company's favor in an earlier author copyright suit. But the same judge wrote that his ruling "does not stand for the proposition that Meta's use of copyrighted materials to train its language models is lawful" — a narrow win that left the door open for exactly the kind of suit filed this week.

The publishers are coming in with a stronger record than the individual authors did. Trade publishers have institutional resources, registered copyrights with statutory damages attached, and textbooks that produce clean, testable infringement examples. A calculus textbook with numbered exercises is easier to match against model output than a literary novel, and that evidentiary advantage runs through the entire complaint.

The Anthropic comparison hangs over the case. A federal judge ruled that training Claude on legally purchased books qualified as fair use, but allowed a class action to proceed over the millions of works Anthropic allegedly pirated. Anthropic settled that suit last year for $1.5 billion, the largest publishing-industry payout in the AI era and a number every plaintiff's lawyer in this space now uses as an anchor.

Educational publishers face a particular threat from generative models. McGraw Hill, Cengage, and Elsevier sell textbooks and journal access into a market where students already turn to chatbots for homework help. If Llama can reproduce sections of Stewart's Calculus on demand, the substitution risk is direct and measurable, not theoretical. That makes the damages theory cleaner than it was for novelists arguing about style.

Related · from this week
Google restricts Meta's access to Gemini AI models
Jaeden Schafer · 4 min read →

The publishers are also asking the court to compel disclosure of every copyrighted work used to train Llama. That request, if granted, would set a discovery precedent every other AI lab would have to reckon with. Training-data opacity has been the industry's default posture, and a court order forcing a full inventory would change the negotiating dynamics for every pending and future copyright suit.

Meta will likely argue, as it has before, that ingestion is transformative and that intermediate copying for training falls within fair use. The verbatim-output claims complicate that defense. Fair use analysis weighs the effect on the market for the original work, and a model that reproduces textbook passages on prompt is, by the publishers' framing, a direct market substitute rather than a transformative use.

The bigger question is whether the industry's settled assumption — that scraped web data is fair game until proven otherwise — survives a sustained legal assault from rightsholders with real budgets. Anthropic's $1.5 billion settlement suggested the answer is no, at least not for pirated sources. Meta's case will test whether that lesson generalizes, and whether the next round of frontier models gets trained on licensed data, synthetic data, or whatever survives the discovery process.

ShareXLinkedInEmail
AI Box

Every AI model. One chat.

The latest models from ChatGPT, Claude, Gemini, Sora, ElevenLabs — 80+ models in a single chat. Compare answers side by side. Pick the best one every time.

  • ChatGPT, Claude, Gemini, Grok, DeepSeek — in one chat
  • Generate images & video with Sora, Veo, Ideogram
  • Compare any two models side by side
  • From $8.99/mo · 80+ models, all included
Try AI Boxaibox.ai
Trusted by 3,000+ teams
Got a tip?

Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.

Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.

AI Box Daily briefingFree · Daily · No fluff

Stay ahead of everyone in AI.

The tightly edited AI news email engineers, founders, and investors actually open. One email. Every weekday. Five minutes to finish.

Loved by 10,000+ AI professionals
Free forever. Unsubscribe with one click.

The briefing read inside teams at

Keep reading

More from Security

Google logo
Business

Google restricts Meta's access to Gemini AI models

The curb signals Google treating Meta as a direct AI rival rather than a customer, per a Financial Times report.

Jaeden Schafer4 min read
Meta logo
Security

Meta reworks AI prompt suggestions after chatbot probes user's children

A viral video showed Meta AI asking a mother to identify her child, then surfacing a photo she says she deleted years ago.

Jaeden Schafer4 min read
Meta logo
Security

Meta rolls back Instagram AI tagging in 3 days after opt-out backlash

The feature let users generate images of public Instagram accounts by default. It survived 72 hours before Meta pulled it.

Jaeden Schafer5 min read