Cloudflare Is Blocking AI Crawlers by Default. What It Means for Your Content's AI Visibility

Cloudflare now blocks mixed-use AI crawlers from ad-supported publisher pages unless the publisher opts in, and plans to charge when AI systems generate value from content. A read for content-driven B2B marketing teams.

  • SEO
  • AI Strategy
  • Content Marketing

Cloudflare made a quiet but significant move this quarter: it began blocking mixed-use crawlers, the kind that combine traditional search indexing, AI agent access, and AI model training, from advertising-supported pages by default, unless the publisher specifically chooses to allow it. Cloudflare has also signaled plans to expand publisher monetization beyond charging for crawling access to charging when AI systems actually generate value from published content.

The stated goal is to push AI companies toward separating their search crawlers from their training and agent crawlers, giving publishers more control over how their content gets used.

Why marketing teams, not just publishers, should care

This is not just a publishing-industry story. It directly affects a question a lot of B2B marketing teams have not thought through: is your own website’s content actually reachable by the AI systems your buyers are asking questions inside of.

If your site sits behind Cloudflare’s protection (a large share of the web does) and the default settings block AI crawlers, your carefully built GEO and AEO strategy may be running into a wall you did not know existed. Content you have optimized to be citation-worthy inside ChatGPT, Perplexity, or Google’s AI Overviews cannot get cited if the crawler that would read it is blocked before it arrives.

Why “mixed-use” crawlers are the crux of the problem

The technical detail that makes this change tricky is that a lot of AI crawlers historically bundled multiple purposes into a single crawler identity — the same bot might be indexing your page for a search-style citation, feeding it into an answer-engine response, and adding it to a training dataset, all under one user-agent string. That made it functionally impossible for a publisher to allow one behavior while blocking another; it was all-or-nothing. Cloudflare’s move is explicitly trying to force AI companies to split those functions into separately identifiable crawlers, which is a prerequisite for the more nuanced control described below — you can’t selectively allow “cite me” while blocking “train on me” if both requests arrive from the same bot with no way to tell them apart. Whether AI companies fully comply with cleanly separated crawlers is still playing out, which is part of why this is worth rechecking periodically rather than configuring once.

Where to actually go check your settings

If you’re on Cloudflare, the relevant controls live under Security → Bots, where you can review and adjust AI crawler behavior per bot category rather than as a single on/off switch. The practical audit: pull up that panel, note which specific bots are currently allowed versus blocked, and cross-reference against the list of AI systems your buyers are actually likely to be using for research — ChatGPT’s crawler, Perplexity’s, Google’s AI Overview infrastructure, Anthropic’s. A default configuration inherited from a security-focused setup (block everything unfamiliar) can end up silently blocking the exact traffic your GEO strategy depends on, and because it’s a crawler-access issue rather than a content issue, nothing about your actual pages will look wrong when you check them manually — the block happens before your content is ever read.

What to check this week

Confirm your own site’s crawler settings. If you use Cloudflare or a similar service, verify explicitly whether AI crawlers are allowed to access your public content, particularly the pages you have invested in making citation-worthy. Do not assume a setting configured months ago still matches your current AI visibility strategy.

Separate the training-data question from the citation question. You may have legitimate reasons to block AI training crawlers (protecting proprietary content from being used to train a competitor’s model) while still wanting AI answer engines to be able to cite and drive brand awareness through your public marketing content. Those are different crawler behaviors, and increasingly, different settings.

Watch this space for more monetization structures. As more infrastructure providers build tools for publishers to charge AI systems for content access, the economics of being findable inside AI answers are likely to keep shifting. What is free and open today may not stay that way, and it is worth revisiting your crawler settings every quarter rather than treating it as a one-time configuration.