Patreon is now actively blocking AI training bots from scraping creator content, extending its partnership with Cloudflare to enforce the policy at the network layer rather than relying on the honor system. The company said Thursday that during testing, individual crawlers' weekly attempts to reach Patreon fell from thousands to zero once the new controls were switched on. That gap — thousands versus zero — is the entire story: the scrapers were reaching the site regardless of what Patreon's robots.txt file asked them to do.
The shift moves Patreon from a request-based model to an enforcement-based one. Since 2023, the platform has used robots.txt directives and a paywall to keep AI crawlers away from paid creator posts. Those measures assumed the crawlers would honor the file. The new numbers suggest many did not.
Patreon is using Cloudflare's AI Crawl Control, which sits in front of the site and identifies training bots by signature and behavior before they can pull content. The company will still allow crawlers that index pages for the purpose of routing users back to Patreon — search discovery is preserved. What is blocked is the scraping-for-training pipeline that feeds foundation models.
Key facts
- 01Patreon is now actively blocking AI training bots via Cloudflare's AI Crawl Control, rather than relying on robots.txt requests.
- 02During testing, individual AI crawlers' weekly attempts to access Patreon dropped from thousands to zero.
- 03Patreon first deployed AI scraping deterrence in 2023, but says crawlers grew more sophisticated and began ignoring robots.txt.
- 04Cloudflare earlier this month began blocking 'mixed-use' crawlers by default on pages that host ads.
- 05Bots that index pages to send users back to Patreon remain allowed; only training scrapers are blocked.
The timing tracks with broader changes to Patreon's surface area. The company recently launched a redesigned Home Feed and a short-post format called Quips, both of which expose more creator content outside the paywall. More discoverable content is also more scrapeable content, which raised the stakes on enforcement.
Cloudflare has been building the infrastructure for this kind of push. The company rolled out Pay Per Crawl, a marketplace that lets publishers charge AI bots for access, and earlier this month changed its default policy so that mixed-use crawlers — bots that both index for search and train models — are blocked by default on pages carrying ads. Patreon is one of the first large creator platforms to adopt the training-specific enforcement layer.
Patreon framed the move as a consent question rather than a compensation one. The company is not announcing a licensing marketplace or a pay-to-train tier. It is drawing a line at unauthorized ingestion, full stop, and leaving the question of paid access for later.
The enforcement gap has become the central problem in AI training data. Robots.txt is a 1994 protocol that assumes goodwill; foundation-model training in 2026 does not reliably provide it. Publishers ranging from news organizations to Reddit have moved to either license their data or block scrapers outright, and lawsuits from The New York Times and others are still working through the courts. Patreon's numbers are one of the cleaner public data points on how much traffic robots.txt was actually stopping — which is to say, in some cases, none.
The counterweight for AI labs is real. Training-data supply is tightening across the open web at the exact moment model developers need more of it to scale next-generation systems. Every major platform that flips from passive to active blocking narrows the available corpus and pushes labs toward licensed deals, synthetic data, or smaller high-quality datasets. Whether that produces better models or just more expensive ones is an open question, and Patreon alone will not settle it.
The precedent matters more than the volume. Patreon hosts a large chunk of the paid creator economy — writers, podcasters, illustrators, video makers — whose work is exactly the kind of long-form, human-generated material that model developers most want. If Cloudflare's AI Crawl Control becomes the default block for creator platforms the way robots.txt was the default request, the balance of power on training data shifts from crawlers to hosts. That is the actual market change here, not the individual policy update.
Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.
Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.




