Skip to main content
Live
Main content

Artists press courts on AI training data, and start winning

Anthropic's $1.5B settlement is the largest copyright payout ever; suits against Google, Meta, Stability, and Suno keep multiplying.

Jaeden Schafer
Editor in Chief · · 5 min read
Anthropic logo

Artists suing AI companies over training data are starting to win. Anthropic agreed to a $1.5 billion settlement — the largest copyright payout on record — after a federal court found it had used pirated ebooks to train Claude, and the company agreed to destroy the pirated trove. Parallel suits against Google, Meta, Stability, Midjourney, Runway, DeviantArt, and the music generator Suno are grinding through the courts, some 5 to 6 years into the effort authors and illustrators say went into the works now used as training fuel.

The Anthropic outcome, Bartz v. Anthropic, is the clearest artist win to date, but it split down the middle. Judge William Alsup ruled that Anthropic violated copyright law by training on pirated ebooks downloaded from the internet — hence the $1.5 billion — but held that a separate trove of secondhand physical books the company bought and scanned under Project Panama qualified as fair use because the training was 'quintessentially transformative.' That transformative-use finding is the ruling every other AI defendant is now leaning on.

Andrea Bartz, the novelist behind We Were Never Here and The Spare Room and the lead plaintiff in the Susman Godfrey case against Anthropic, disputes the fair-use half of the split. She argues that even a library cannot legally buy a physical copy of a book, scan it, and lend it out as an ebook, and hopes higher courts revisit the standard. Her broader complaint is structural: that entire libraries were ingested without permission or payment as a foundation for commercial products.

I felt violated, shocked, alarmed
Andrea Bartz, Novelist and lead plaintiff in Bartz v. Anthropic

Key facts

  • 01Anthropic agreed to a $1.5 billion settlement — the largest copyright payout on record — over pirated ebooks used to train Claude.
  • 02A class action by Sarah Andersen, Karla Ortiz, and Kelly McKernan against Stability, Midjourney, DeviantArt, and Runway has been in court since January 2023.
  • 03Judge William Alsup ruled Anthropic's use of legally purchased scanned books was 'quintessentially transformative' fair use, splitting the ruling.
  • 04Musician Sam Kogon is suing Google over Lyria and ProducerAI, alleging YouTube's terms of service don't authorize training on uploaded content.
  • 05In Kadrey v. Meta, the judge dismissed many claims for lack of market harm; narrower copyright and pirated-materials claims continue.

The visual-arts front opened first. Sarah Andersen, whose Sarah's Scribbles webcomic she describes as a 'complex culmination of my education, the comics I devoured as a child, and the many small choices that make up the sum of my life,' filed suit alongside Karla Ortiz, Kelly McKernan, and other illustrators in January 2023 against Stability, Midjourney, DeviantArt, and Runway AI. That was two months after ChatGPT's November 2022 debut, when generative AI was still a curiosity rather than a national-security concern. The case has crawled through the courts ever since.

Author Kirk Wallace Johnson, who wrote The Feather Thief and The Fishermen and the Dragon, found his books in the searchable training-data set The Atlantic published and proactively contacted Susman Godfrey to join the fight. His argument is that celebrity plaintiffs distort the debate. 'AI could never write The Godfather,' he said, 'But AI could write a mediocre film. AI could write a mediocre book. And there are tons of authors and screenwriters that live in that space.' The economic risk, in his framing, falls on the working middle of the creative industry — not on the Agatha Christies.

The music suits attack from a different angle. Sam Kogon, an independent musician, is the lead plaintiff in an ongoing case against Google's Lyria and ProducerAI music engines. Rather than pure copyright, Kogon's lawyers argue Google improperly used its Content ID system and YouTube uploads to train the models, in violation of the platform's terms of service. Google has filed a motion to dismiss, arguing the YouTube TOS grants it broad rights to reproduce, distribute, and prepare derivative works from uploaded content.

Krystle Delgado, an entertainment and IP lawyer who runs the Top Music Attorney channel, said most creators do not read the YouTube terms as a license to remake their content, but that the fine print grants YouTube an 'irrevocable perpetual license.' Google spokesperson Jack Malon told The Verge the company has said for several years it uses uploaded content to improve product experiences across YouTube and Google, including for machine learning and AI applications.

Meta is fighting a separate suit brought by Richard Kadrey, Sarah Silverman, Christopher Golden, and others over the use of their books to train Llama. The judge in Kadrey v. Meta dismissed many of the initial claims for lack of demonstrated market harm, but a narrower set focused on direct copyright infringement and the use of pirated materials is still active. The plaintiffs' complaint argued that AI-generated books 'probably wouldn't have much of an effect on the market for the works of Agatha Christie,' but 'could very well prevent the next Agatha Christie from getting noticed or selling enough books to keep writing.'

Related · from this week
Authors accuse publishers and agents of overreaching on Anthropic settlement payouts
Jaeden Schafer · 5 min read →

The defense playbook across all these cases is largely the same: training is transformative, terms of service authorized the ingestion, and market harm is speculative. That defense has already partially failed for Anthropic on the pirated-books question and is under active challenge on the TOS question with Google. Delgado said the direction of travel — and the shift in public sentiment — is starting to favor plaintiffs.

The courts and the judges seem to be starting to lean our way, and the court of public opinion too.
Krystle Delgado, Entertainment and IP lawyer

The commercial stakes are the real story. A precedent that training on pirated material is straightforwardly infringing would force every major lab to audit its data provenance and negotiate licensing at scale, on top of the compute bills already stretching capex budgets. A precedent that legally acquired scans qualify as fair use, meanwhile, effectively legalizes the current book-training pipeline for anyone willing to buy and scan. The Anthropic settlement suggests the industry can absorb ten-figure payouts as a cost of doing business — which means the lasting question is not whether artists collect, but whether courts eventually force the pipelines themselves to change.

ShareXLinkedInEmail
AI Box

Every AI model. One chat.

The latest models from ChatGPT, Claude, Gemini, Sora, ElevenLabs — 80+ models in a single chat. Compare answers side by side. Pick the best one every time.

  • ChatGPT, Claude, Gemini, Grok, DeepSeek — in one chat
  • Generate images & video with Sora, Veo, Ideogram
  • Compare any two models side by side
  • From $8.99/mo · 80+ models, all included
Try AI Boxaibox.ai
Trusted by 3,000+ teams
Got a tip?

Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.

Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.

AI Box Daily briefingFree · Daily · No fluff

Stay ahead of everyone in AI.

The tightly edited AI news email engineers, founders, and investors actually open. One email. Every weekday. Five minutes to finish.

Loved by 10,000+ AI professionals
Free forever. Unsubscribe with one click.

The briefing read inside teams at

Keep reading

More from Security

Anthropic logo
Security

Authors accuse publishers and agents of overreaching on Anthropic settlement payouts

Writers say HarperCollins and others are claiming shares of the $1.5B fund on books whose rights reverted years ago.

Jaeden Schafer5 min read
OpenAI logo
Security

DOJ backs OpenAI in New York Times copyright case, calls AI training fair use

The Justice Department's statement of interest argues a Times win would 'thwart' American AI leadership and misread fair use doctrine.

Jaeden Schafer5 min read
Google logo
Security

Munich court holds Google liable for false claims in AI Overviews

A German ruling says generative search outputs aren't third-party speech — the operator owns the words, and the disclaimer doesn't save it.

Jaeden Schafer5 min read