Skip to main content
Live
Main content

Publishers sue Google, allege Gemini was trained on pirated books

Hachette, Cengage, Elsevier, and Scott Turow accuse Google of scraping books from Google Books and Google Play to train Gemini without permission.

Jaeden Schafer
Editor in Chief · · 5 min read
Google logo

Hachette, Cengage, Elsevier, author Scott Turow, and the writers' group S.C.R.I.B.E. filed a class action against Google on July 14, 2026, alleging the company trained its Gemini models on copyrighted books it obtained through scope-limited publishing programs. The suit, filed in the U.S. District Court for the Southern District of New York, also accuses Google of stripping or altering copyright management information to hide the source of its training data. The plaintiffs cite an internal Google document that allegedly warned using copyrighted books for AI training could be "highly problematic" and expose the company to "$10Bs-$100Bs in potential fines."

The core allegation turns on the specific relationship publishers have with Google. For years, publishers and authors handed Google copyrighted books under narrow terms — for indexing in Google Books, which returns short snippets and bibliographic data, and for sale through the Google Play store. The complaint argues Google took those same corpora and reused them to train Gemini without a separate license.

That distinction matters. Prior AI copyright cases have hinged on whether scraping from the open web qualifies as fair use. Here, publishers can point to specific contracts governing specific uses, and argue Google exceeded them. The complaint frames the alleged copying as a breach of the scope of programs Google itself designed.

Key facts

  • 01Hachette, Cengage, Elsevier, Scott Turow, and S.C.R.I.B.E. filed a class action against Google in the U.S. District Court for the Southern District of New York on July 14, 2026.
  • 02Plaintiffs allege Google trained Gemini on books submitted under scope-limited programs including Google Books and Google Play.
  • 03The complaint cites an internal Google document warning of '$10Bs-$100Bs in potential fines' from using copyrighted books to train AI.
  • 04Anthropic previously paid $1.5B to settle a similar piracy claim, with about 500,000 writers eligible for at least $3,000 each.
  • 05Two prior California rulings sided with AI companies on fair use, but SDNY gives publishers a fresh venue.

The lawsuit is one of a growing pile against AI developers including Google, Meta, OpenAI, and Anthropic. Two early rulings in California went the AI companies' way on fair use grounds, applying a U.S. copyright statute that predates the modern web. Those decisions are not binding on a New York judge, and the plaintiffs are betting SDNY reads the record differently.

Google illegally copied works from all these scope-limited programs for AI training, knowing it lacked authorization to do so
Scott Turow, Author and named plaintiff

The most instructive precedent is financial rather than doctrinal. Anthropic was fined $1.5 billion for pirating works used to train its models — the largest payout in U.S. copyright history. Roughly 500,000 writers were eligible for payments of at least $3,000 each, though many opted out to pursue their own claims. The internal Google memo's own $10B–$100B range suggests Google's lawyers modeled a substantially larger exposure.

Google did not immediately respond to a request for comment. The company has consistently argued in prior filings that training large language models on lawfully accessed text is transformative and protected under fair use. Whether that argument survives when the "lawful access" came through a contractually limited program is the question SDNY will now weigh.

The plaintiff list is worth reading closely. Hachette, Cengage, and Elsevier cover trade publishing, higher-education textbooks, and scientific journals — three revenue pools that intersect directly with what Gemini and its competitors are being sold to do. Cengage and Elsevier in particular license their catalogs to enterprise customers at premium rates. If a general-purpose chatbot can regurgitate textbook and journal content on demand, the licensing case against AI developers writes itself.

Turow adds an individual-author dimension that mirrors the Anthropic settlement structure. If the case survives motions to dismiss and reaches class certification, the plaintiff pool could balloon quickly. The Anthropic precedent showed how a per-writer floor payment scales: 500,000 claimants at $3,000 apiece is $1.5 billion before any per-work multipliers.

Related · from this week
Google fixes 1,072 Chrome bugs in one month using AI, more than the past two years combined
Jaeden Schafer · 5 min read →

The counterargument Google will make is that snippet-level exposure in Google Books was upheld as fair use in the landmark Authors Guild v. Google decision more than a decade ago. But that ruling addressed search, not generative output. Publishers will argue that training a model that can produce book-length prose is a different act from returning a two-sentence snippet, regardless of whether the underlying corpus is the same.

There is also the alleged tampering with copyright management information. That is a separate statutory claim under the DMCA with its own damages framework, and it does not require the plaintiffs to defeat fair use to collect. If discovery surfaces evidence Google stripped attribution data from ingested books, the case gets harder to settle quietly.

The lawsuit lands at an awkward moment for Google's AI monetization push. Gemini is central to the company's enterprise pitch, and any injunction touching training data or requiring retraining on cleaned corpora would be operationally painful. Even without an injunction, a settlement in the Anthropic range would be absorbable for a company of Google's size, but the internal memo's higher-end $100B figure would not be.

The SDNY filing pushes the AI copyright fight into its most consequential venue yet, with a plaintiff coalition that has both the contract theory and the corporate paper trail to make the case land. Two California rulings gave the AI industry an early narrative win on fair use; a contrary ruling in New York, or a settlement priced anywhere near Google's own internal estimate, would reset licensing economics for every frontier lab shipping a model trained on books.

ShareXLinkedInEmail
AI Box

Every AI model. One chat.

The latest models from ChatGPT, Claude, Gemini, Sora, ElevenLabs — 80+ models in a single chat. Compare answers side by side. Pick the best one every time.

  • ChatGPT, Claude, Gemini, Grok, DeepSeek — in one chat
  • Generate images & video with Sora, Veo, Ideogram
  • Compare any two models side by side
  • From $8.99/mo · 80+ models, all included
Try AI Boxaibox.ai
Trusted by 3,000+ teams
Got a tip?

Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.

Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.

AI Box Daily briefingFree · Daily · No fluff

Stay ahead of everyone in AI.

The tightly edited AI news email engineers, founders, and investors actually open. One email. Every weekday. Five minutes to finish.

Loved by 10,000+ AI professionals
Free forever. Unsubscribe with one click.

The briefing read inside teams at

Keep reading

More from Security

Google logo
Security

Google fixes 1,072 Chrome bugs in one month using AI, more than the past two years combined

Chrome 149 and 150 patched more security flaws than the previous 23 releases combined, as Google credits Gemini for the surge.

Jaeden Schafer5 min read
Google logo
Security

Google launches Gemini 3.5 Flash Cyber to undercut Anthropic's Mythos

The new security model runs at a fraction of Mythos 5's cost and found 55 confirmed bugs in V8, beating Claude Opus 4.6's 36.

Jaeden Schafer4 min read
Google logo
Security

Google sues Chinese network Outsider Enterprise over Gemini-built scam sites

The group allegedly used Gemini to spin up 9,000 fake sites and blast 2.5 million scam texts to Android users.

Jaeden Schafer5 min read