Skip to main content
Live
Main content

Patreon starts blocking AI scrapers with Cloudflare, dropping robots.txt approach

The membership platform is shifting from polite requests to active enforcement, cutting weekly scrape attempts from thousands to zero.

Jaeden Schafer
Editor in Chief · · 4 min read
Patreon starts blocking AI scrapers with Cloudflare, dropping robots.txt approach

Patreon is now actively blocking AI training bots from scraping creator content, extending its partnership with Cloudflare to enforce the policy at the network layer rather than relying on the honor system. The company said Thursday that during testing, individual crawlers' weekly attempts to reach Patreon fell from thousands to zero once the new controls were switched on. That gap — thousands versus zero — is the entire story: the scrapers were reaching the site regardless of what Patreon's robots.txt file asked them to do.

The shift moves Patreon from a request-based model to an enforcement-based one. Since 2023, the platform has used robots.txt directives and a paywall to keep AI crawlers away from paid creator posts. Those measures assumed the crawlers would honor the file. The new numbers suggest many did not.

Patreon is using Cloudflare's AI Crawl Control, which sits in front of the site and identifies training bots by signature and behavior before they can pull content. The company will still allow crawlers that index pages for the purpose of routing users back to Patreon — search discovery is preserved. What is blocked is the scraping-for-training pipeline that feeds foundation models.

Key facts

  • 01Patreon is now actively blocking AI training bots via Cloudflare's AI Crawl Control, rather than relying on robots.txt requests.
  • 02During testing, individual AI crawlers' weekly attempts to access Patreon dropped from thousands to zero.
  • 03Patreon first deployed AI scraping deterrence in 2023, but says crawlers grew more sophisticated and began ignoring robots.txt.
  • 04Cloudflare earlier this month began blocking 'mixed-use' crawlers by default on pages that host ads.
  • 05Bots that index pages to send users back to Patreon remain allowed; only training scrapers are blocked.

The timing tracks with broader changes to Patreon's surface area. The company recently launched a redesigned Home Feed and a short-post format called Quips, both of which expose more creator content outside the paywall. More discoverable content is also more scrapeable content, which raised the stakes on enforcement.

Cloudflare has been building the infrastructure for this kind of push. The company rolled out Pay Per Crawl, a marketplace that lets publishers charge AI bots for access, and earlier this month changed its default policy so that mixed-use crawlers — bots that both index for search and train models — are blocked by default on pages carrying ads. Patreon is one of the first large creator platforms to adopt the training-specific enforcement layer.

Patreon framed the move as a consent question rather than a compensation one. The company is not announcing a licensing marketplace or a pay-to-train tier. It is drawing a line at unauthorized ingestion, full stop, and leaving the question of paid access for later.

The enforcement gap has become the central problem in AI training data. Robots.txt is a 1994 protocol that assumes goodwill; foundation-model training in 2026 does not reliably provide it. Publishers ranging from news organizations to Reddit have moved to either license their data or block scrapers outright, and lawsuits from The New York Times and others are still working through the courts. Patreon's numbers are one of the cleaner public data points on how much traffic robots.txt was actually stopping — which is to say, in some cases, none.

The counterweight for AI labs is real. Training-data supply is tightening across the open web at the exact moment model developers need more of it to scale next-generation systems. Every major platform that flips from passive to active blocking narrows the available corpus and pushes labs toward licensed deals, synthetic data, or smaller high-quality datasets. Whether that produces better models or just more expensive ones is an open question, and Patreon alone will not settle it.

Related · from this week
Cloudflare will block mixed-use AI crawlers by default from September 2026
Jaeden Schafer · 5 min read →

The precedent matters more than the volume. Patreon hosts a large chunk of the paid creator economy — writers, podcasters, illustrators, video makers — whose work is exactly the kind of long-form, human-generated material that model developers most want. If Cloudflare's AI Crawl Control becomes the default block for creator platforms the way robots.txt was the default request, the balance of power on training data shifts from crawlers to hosts. That is the actual market change here, not the individual policy update.

ShareXLinkedInEmail
AI Box

Every AI model. One chat.

The latest models from ChatGPT, Claude, Gemini, Sora, ElevenLabs — 80+ models in a single chat. Compare answers side by side. Pick the best one every time.

  • ChatGPT, Claude, Gemini, Grok, DeepSeek — in one chat
  • Generate images & video with Sora, Veo, Ideogram
  • Compare any two models side by side
  • From $8.99/mo · 80+ models, all included
Try AI Boxaibox.ai
Trusted by 3,000+ teams
Got a tip?

Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.

Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.

AI Box Daily briefingFree · Daily · No fluff

Stay ahead of everyone in AI.

The tightly edited AI news email engineers, founders, and investors actually open. One email. Every weekday. Five minutes to finish.

Loved by 10,000+ AI professionals
Free forever. Unsubscribe with one click.

The briefing read inside teams at

Keep reading

More from Security

Cloudflare will block mixed-use AI crawlers by default from September 2026
Business

Cloudflare will block mixed-use AI crawlers by default from September 2026

New defaults force AI companies to separate search bots from training and agent crawlers, or lose access to ad-supported pages.

Jaeden Schafer5 min read
ShieldFont poisons AI scrapers by swapping words with ligatures
Security

ShieldFont poisons AI scrapers by swapping words with ligatures

A new font replaces 24.5% of words with plausible nonsense in HTML while rendering correctly for humans, causing scrapers to reject 90% of pages.

Jaeden Schafer5 min read
AWS rebuilds OpenSearch Serverless for agent traffic as bots near half the web
Business

AWS rebuilds OpenSearch Serverless for agent traffic as bots near half the web

The new Serverless decouples compute from storage and scales to zero — Cloudflare says non-human traffic will pass human traffic in early 2027.

Jaeden Schafer5 min read