Skip to main content
Live
Main content

Cloudflare will block mixed-use AI crawlers by default from September 2026

New defaults force AI companies to separate search bots from training and agent crawlers, or lose access to ad-supported pages.

Jaeden Schafer
Editor in Chief · · 5 min read
Cloudflare will block mixed-use AI crawlers by default from September 2026

Cloudflare will block mixed-use AI crawlers by default from any ad-supported page starting September 15, 2026, a policy change that forces AI companies to either separate their search bots from their training and agent crawlers or lose access to a large slice of the open web. The new defaults apply to new Cloudflare customers, any new sites set up by existing customers, and every existing free-tier customer. Publishers who want to keep letting the blended bots through will have to change the setting themselves.

The rule targets crawlers that blend three jobs into one user agent: indexing for traditional search, fetching pages for AI agents, and pulling content for model training. Cloudflare's argument is that publishers largely want to remain discoverable in search but do not want their intellectual property fed into training corpora or answer engines for free. Under the new default, a crawler that refuses to declare which of those three jobs it is doing gets blocked.

Cloudflare singled out the operator of the world's largest search engine, a clear reference to Google, saying it has access to roughly 2x more information than other AI companies because opting out of AI use has historically meant opting out of Search discovery. Google has pushed back on that framing before, pointing to Google Extended, a bot that lets site owners exclude their content from training and from products such as Gemini Apps and Vertex API without affecting Search inclusion. Googlebot itself, however, still crawls for Search features including AI Overviews and AI Mode, which is the specific overlap the new Cloudflare defaults are designed to expose.

Now that the majority of traffic on the Internet is non-human, we must go further and act faster so that a sustainable ecosystem can emerge.
Matthew Prince, Cloudflare co-founder and CEO

Key facts

  • 01Starting September 15, 2026, Cloudflare will block mixed-use AI crawlers by default from any page hosting ads.
  • 02The new defaults apply to new customers, new sites from existing customers, and all existing free-tier customers.
  • 03Cloudflare data shows over 50% of crawl traffic from AI crawlers is spent re-fetching unchanged pages.
  • 04Pay Per Crawl is evolving into Pay Per Use, charging AI companies when publisher content creates value rather than only on fetch.
  • 05Ceramic.ai and You.com are the first two partners implementing the model.

The timing reflects a shift Cloudflare says arrived earlier than expected: bots have already overtaken humans as the majority of internet traffic, a milestone the company had not projected until next year. CEO Matthew Prince framed the deadline as a response to that pace, arguing that a workable content economy needs bots to declare intent rather than hide behind generic user agents.

Alongside the default change, Cloudflare is evolving its Pay Per Crawl marketplace into a broader Pay Per Use model. The original product let publishers charge AI bots per fetch. The new version lets publishers charge when their content generates value downstream, not only when it is scraped. Cloudflare's own data underlines why fetch-based pricing has limits: more than 50% of crawl traffic from AI crawlers is spent re-fetching pages that have not changed, which is expensive for the AI companies and offers little to publishers.

The first two AI companies wired into the new model are Ceramic.ai and You.com. When a publisher opts in, they get paid when their content surfaces in Ceramic's AI search results, or when You.com pulls one of their premium pages. Cloudflare says other AI companies can customize the arrangement to match their own product mechanics.

For the frontier labs, the practical implication is operational rather than existential. OpenAI, Anthropic, Google, and xAI already run dedicated user agents for training and for retrieval, and any lab that wants to keep crawling ad-supported pages behind Cloudflare can do so by publishing clearly separated crawlers with declared purposes. The friction lands hardest on smaller AI companies and on any operator that has been letting a single blended bot handle everything, because the September deadline forces engineering work that was previously optional.

Publishers get leverage they have not had before. Cloudflare sits in front of a large share of the web, and a default-on block is qualitatively different from an opt-in tool that most site owners never configure. The company has spent the past two years building crawler-control features aimed at exactly this moment, and the shift from Pay Per Crawl to Pay Per Use is the clearest sign yet that the pricing debate is moving from access to outcomes.

Related · from this week
Anthropic acquires Stainless for over $300M, pulls SDK tool from OpenAI and Google
Jaeden Schafer · 5 min read →

The counterweight is enforcement and coverage. Cloudflare's defaults apply only to sites on Cloudflare, and a determined crawler can still route around any specific network. The Pay Per Use economics also depend on AI companies reporting downstream value honestly, which is a harder telemetry problem than logging a fetch. And publishers on paid Cloudflare plans who have already customized their crawler settings will not see the new defaults applied automatically, which limits the immediate blast radius.

The bigger market signal is that the era of ambiguous crawling is closing. AI companies that want durable access to the open web will need to publish separate, purpose-declared bots and negotiate commercial terms for training and answer-engine use, rather than relying on the assumption that a single search-flavored user agent buys them everything. That reshapes the cost structure for anyone building retrieval-heavy AI products, and it hands publishers a pricing surface they can actually meter.

ShareXLinkedInEmail
AI Box

Every AI model. One chat.

The latest models from ChatGPT, Claude, Gemini, Sora, ElevenLabs — 80+ models in a single chat. Compare answers side by side. Pick the best one every time.

  • ChatGPT, Claude, Gemini, Grok, DeepSeek — in one chat
  • Generate images & video with Sora, Veo, Ideogram
  • Compare any two models side by side
  • From $8.99/mo · 80+ models, all included
Try AI Boxaibox.ai
Trusted by 3,000+ teams
Got a tip?

Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.

Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.

AI Box Daily briefingFree · Daily · No fluff

Stay ahead of everyone in AI.

The tightly edited AI news email engineers, founders, and investors actually open. One email. Every weekday. Five minutes to finish.

Loved by 10,000+ AI professionals
Free forever. Unsubscribe with one click.

The briefing read inside teams at

Keep reading

More from Business

Anthropic logo
Business

Anthropic acquires Stainless for over $300M, pulls SDK tool from OpenAI and Google

The deal removes a key infrastructure supplier from rivals' hands; Anthropic will wind down all hosted Stainless products.

Jaeden Schafer5 min read
Patreon starts blocking AI scrapers with Cloudflare, dropping robots.txt approach
Security

Patreon starts blocking AI scrapers with Cloudflare, dropping robots.txt approach

The membership platform is shifting from polite requests to active enforcement, cutting weekly scrape attempts from thousands to zero.

Jaeden Schafer4 min read
Microsoft logo
Business

Microsoft's carbon emissions rose 25% in 2025 as AI data centers expanded

The company's 2026 sustainability report says emissions hit 34M metric tons and warns sustainability solutions can't keep pace with AI demand.

Jaeden Schafer4 min read