Cara, the art portfolio app built by photographer Jingna Zhang for the 1.5 million artists who refuse to let their work train AI models, was scraped three times in ten days this month. The first hit, on August 13, pulled a 12-terabyte archive of 12 million works — effectively the platform's entire public library — for a total compute cost of under $10. The scraper posted the trove on the subreddit r/DefendingAIArt and called it a fun project.
The second scrape lifted 8.5 million links plus usernames, titles, and tags, and uploaded them to Hugging Face under the handle CaptiveDreamer. Hugging Face agreed to remove the personal metadata but declined to take down the URLs, telling takedown filers that no copies of the artworks are hosted here and that further copyright reports on the same basis will not change this outcome. On August 22, a third scraper pulled 123,000 images along with user bios and posted them to Academic Torrents.
Zhang launched a GoFundMe with a $120,000 target to cover legal fees and had raised more than $100,000 by Thursday. She is separately a plaintiff in two ongoing class actions brought by visual artists, one against Stability AI, Midjourney, and others, the second against Google, both alleging copyrighted images were used to train commercial image generators.
Key facts
- 01A scraper pulled a 12-terabyte archive of 12 million works from Cara — roughly its entire public library — for less than $10.
- 02Cara hosts about 1.5 million artists who joined specifically to avoid AI training scrapes.
- 03A second scraper uploaded 8.5 million links to Hugging Face; a third grabbed 123,000 images and posted them to Academic Torrents on August 22.
- 04Zhang's GoFundMe for legal fees has raised more than $100,000 of a $120,000 goal as of Thursday.
- 05The first scraper, 'Heft,' has apologized, deleted the dataset, and is co-building Lantern, an open-source tool that scans new AI datasets for artists' fingerprinted work.
Then something unusual happened. The person who kicked off the scraping wave apologized, deleted his dataset, and joined Cara's Discord as a volunteer. He goes by Heft — a screen name he asked Wired to use after receiving doxing and death threats — and describes himself as a North American student with a background in software and an interest in digital preservation.
Heft told Wired he originally had no plan to publish the data. He made what he called a foolish decision to attempt to ragebait with the dataset on Reddit and was carried away by trolling in the comments. Watching artists in the aftermath — some posting about panic attacks, others deleting entire portfolios from the internet — changed his read of the situation.
His technical assessment of the scrape's actual training risk is worth noting: 12 million images is not a lot to train an image model, most commercial AI labs train from large-scale web scrapes available through open-data organizations like LAION, and he considers it highly unlikely that OpenAI or Anthropic is scanning every new Hugging Face dataset to train on. In other words, the direct model-training threat from these dumps is smaller than the community reaction suggests. The harm was real — but it was harassment and loss of control, not a Stability-scale training pipeline.
“In retrospect, not only deliberately targeting Cara but presenting it the way I did in the post was cruel and thoughtless”— Heft, the scraper, now Cara collaborator
Zhang is candid about the limits of what Cara can do. The site already filters AI-generated images and offers Glaze, the University of Chicago tool that perturbs images to disrupt style mimicry. New login gates went up as a temporary measure. But she is blunt that no site can be sealed off, and she has resisted framing Cara as a safe harbor.
Some users have deleted their work and left the platform anyway, which Zhang says she supports even as she pushes back on the premise. Bigger platforms scrape more; Instagram content is explicitly training data for Meta under its terms. Cara at least forbids it — a point Heft, who insists no site can be made truly unscrapable, agrees with in principle even as he helps Zhang plug holes.
The tool the two are now building is called Lantern. It lets artists generate a one-way fingerprint of an image without Cara storing the image itself, then regularly scans new publicly available AI image datasets. When an artist's fingerprint matches something in a fresh dataset, the artist gets a notification and a link so they can file a removal request or a takedown notice. It's a detection layer, not a prevention layer — a bet that surveillance of the datasets is more tractable than defense of the source.
Lantern is open-source and early. It will not stop the next scrape, and it depends on scrapers continuing to publish their trophies on Hugging Face, Academic Torrents, and similar venues. Heft is clear-eyed that determined actors can defeat almost any Cara-side control in minutes. The tool is a workaround for a regulatory vacuum, not a fix for it.
That vacuum is the real story here. Copyright law, as Zhang notes, has not caught up to bulk data harvests, and Hugging Face's response to the CaptiveDreamer dataset — that hosting URLs pointing to public images is not a copyright violation the platform will act on — is the current legal posture across most dataset hosts. The artist class actions against Stability AI, Midjourney, and Google are still working through court; none have produced a binding precedent that would let Cara issue effective takedowns.
The Cara episode compresses the artist-versus-AI fight into a clean case study: an opt-out community, a $10 attack that vacuumed the entire library, a host that will not remove the links, and a defender tool that can only tell you after the fact that your work is in a dataset. Lantern is a reasonable response to a broken enforcement regime, and the fact that it exists because the scraper switched sides is the most interesting product-development story of the month. But detection tools do not solve the underlying problem, which is that public-image hosting and dataset publication currently sit in the same unregulated space. Until a court or a legislature draws a line — on scraping, on dataset republication, or on training-data provenance — every Cara-like platform is running the same experiment, and the artists are the control group.
Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.
Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.




