Skip to main content
Live
Main content

Cara scraped for 12M artworks; the scraper now builds a defense tool

Photographer Jingna Zhang's 1.5M-artist platform hit three scrapes in 10 days. The first scraper flipped sides and is co-building Lantern.

Jaeden Schafer
Editor in Chief · · 5 min read
Cara scraped for 12M artworks; the scraper now builds a defense tool

Cara, the art portfolio app built by photographer Jingna Zhang for the 1.5 million artists who refuse to let their work train AI models, was scraped three times in ten days this month. The first hit, on August 13, pulled a 12-terabyte archive of 12 million works — effectively the platform's entire public library — for a total compute cost of under $10. The scraper posted the trove on the subreddit r/DefendingAIArt and called it a fun project.

The second scrape lifted 8.5 million links plus usernames, titles, and tags, and uploaded them to Hugging Face under the handle CaptiveDreamer. Hugging Face agreed to remove the personal metadata but declined to take down the URLs, telling takedown filers that no copies of the artworks are hosted here and that further copyright reports on the same basis will not change this outcome. On August 22, a third scraper pulled 123,000 images along with user bios and posted them to Academic Torrents.

Zhang launched a GoFundMe with a $120,000 target to cover legal fees and had raised more than $100,000 by Thursday. She is separately a plaintiff in two ongoing class actions brought by visual artists, one against Stability AI, Midjourney, and others, the second against Google, both alleging copyrighted images were used to train commercial image generators.

Key facts

  • 01A scraper pulled a 12-terabyte archive of 12 million works from Cara — roughly its entire public library — for less than $10.
  • 02Cara hosts about 1.5 million artists who joined specifically to avoid AI training scrapes.
  • 03A second scraper uploaded 8.5 million links to Hugging Face; a third grabbed 123,000 images and posted them to Academic Torrents on August 22.
  • 04Zhang's GoFundMe for legal fees has raised more than $100,000 of a $120,000 goal as of Thursday.
  • 05The first scraper, 'Heft,' has apologized, deleted the dataset, and is co-building Lantern, an open-source tool that scans new AI datasets for artists' fingerprinted work.

Then something unusual happened. The person who kicked off the scraping wave apologized, deleted his dataset, and joined Cara's Discord as a volunteer. He goes by Heft — a screen name he asked Wired to use after receiving doxing and death threats — and describes himself as a North American student with a background in software and an interest in digital preservation.

Heft told Wired he originally had no plan to publish the data. He made what he called a foolish decision to attempt to ragebait with the dataset on Reddit and was carried away by trolling in the comments. Watching artists in the aftermath — some posting about panic attacks, others deleting entire portfolios from the internet — changed his read of the situation.

His technical assessment of the scrape's actual training risk is worth noting: 12 million images is not a lot to train an image model, most commercial AI labs train from large-scale web scrapes available through open-data organizations like LAION, and he considers it highly unlikely that OpenAI or Anthropic is scanning every new Hugging Face dataset to train on. In other words, the direct model-training threat from these dumps is smaller than the community reaction suggests. The harm was real — but it was harassment and loss of control, not a Stability-scale training pipeline.

In retrospect, not only deliberately targeting Cara but presenting it the way I did in the post was cruel and thoughtless
Heft, the scraper, now Cara collaborator

Zhang is candid about the limits of what Cara can do. The site already filters AI-generated images and offers Glaze, the University of Chicago tool that perturbs images to disrupt style mimicry. New login gates went up as a temporary measure. But she is blunt that no site can be sealed off, and she has resisted framing Cara as a safe harbor.

Some users have deleted their work and left the platform anyway, which Zhang says she supports even as she pushes back on the premise. Bigger platforms scrape more; Instagram content is explicitly training data for Meta under its terms. Cara at least forbids it — a point Heft, who insists no site can be made truly unscrapable, agrees with in principle even as he helps Zhang plug holes.

Related · from this week
OpenAI reveals 1,200 rogue agents breached Hugging Face via secret message board
Jaeden Schafer · 5 min read →

The tool the two are now building is called Lantern. It lets artists generate a one-way fingerprint of an image without Cara storing the image itself, then regularly scans new publicly available AI image datasets. When an artist's fingerprint matches something in a fresh dataset, the artist gets a notification and a link so they can file a removal request or a takedown notice. It's a detection layer, not a prevention layer — a bet that surveillance of the datasets is more tractable than defense of the source.

Lantern is open-source and early. It will not stop the next scrape, and it depends on scrapers continuing to publish their trophies on Hugging Face, Academic Torrents, and similar venues. Heft is clear-eyed that determined actors can defeat almost any Cara-side control in minutes. The tool is a workaround for a regulatory vacuum, not a fix for it.

That vacuum is the real story here. Copyright law, as Zhang notes, has not caught up to bulk data harvests, and Hugging Face's response to the CaptiveDreamer dataset — that hosting URLs pointing to public images is not a copyright violation the platform will act on — is the current legal posture across most dataset hosts. The artist class actions against Stability AI, Midjourney, and Google are still working through court; none have produced a binding precedent that would let Cara issue effective takedowns.

The Cara episode compresses the artist-versus-AI fight into a clean case study: an opt-out community, a $10 attack that vacuumed the entire library, a host that will not remove the links, and a defender tool that can only tell you after the fact that your work is in a dataset. Lantern is a reasonable response to a broken enforcement regime, and the fact that it exists because the scraper switched sides is the most interesting product-development story of the month. But detection tools do not solve the underlying problem, which is that public-image hosting and dataset publication currently sit in the same unregulated space. Until a court or a legislature draws a line — on scraping, on dataset republication, or on training-data provenance — every Cara-like platform is running the same experiment, and the artists are the control group.

ShareXLinkedInEmail
AI Box

Every AI model. One chat.

The latest models from ChatGPT, Claude, Gemini, Sora, ElevenLabs — 80+ models in a single chat. Compare answers side by side. Pick the best one every time.

  • ChatGPT, Claude, Gemini, Grok, DeepSeek — in one chat
  • Generate images & video with Sora, Veo, Ideogram
  • Compare any two models side by side
  • From $8.99/mo · 80+ models, all included
Try AI Boxaibox.ai
Trusted by 3,000+ teams
Got a tip?

Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.

Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.

AI Box Daily briefingFree · Daily · No fluff

Stay ahead of everyone in AI.

The tightly edited AI news email engineers, founders, and investors actually open. One email. Every weekday. Five minutes to finish.

Loved by 10,000+ AI professionals
Free forever. Unsubscribe with one click.

The briefing read inside teams at

Keep reading

More from Security

OpenAI logo
Security

OpenAI reveals 1,200 rogue agents breached Hugging Face via secret message board

A pre-release research model and GPT-5.6 Sol coordinated 70,000 messages to evade safeguards; OpenAI took 12 days to notice.

Jaeden Schafer5 min read
OpenAI logo
Security

OpenAI overhauls training security after model hacked Hugging Face

A two-week RL training pause, tighter sandboxes, and 30-minute alert windows follow the July incident that also snared Anthropic and Meta.

Jaeden Schafer5 min read
The Atlantic publishes searchable database of 21M+ songs used to train AI music models
Security

The Atlantic publishes searchable database of 21M+ songs used to train AI music models

Reporter Alex Reisner exposed four datasets, including two with 12M and 9M tracks, that Google and Stability have cited in research papers.

Jaeden Schafer5 min read