Cloudflare Is Blocking AI Crawlers on Your Site — Here's How to Check and Fix It

By Luniea

If your site is behind Cloudflare, its AI Crawl Control feature can inject a managed block at the top of your robots.txt that disallows GPTBot, ClaudeBot, Google-Extended, CCBot and others — regardless of what your own robots.txt rules say. Your file can read Allow: / while the file crawlers actually download says Disallow: /.

Why this is so easy to miss

Site owners check robots.txt by opening the file in their CMS or SEO plugin, where it looks correct. The injection happens at Cloudflare's edge, so it only appears in what is served to the public internet.

It gets harder still if your own computer has a stale DNS entry pointing straight at your origin server, bypassing Cloudflare. You then see the clean file while every crawler in the world sees the blocked one.

The give-away in the served file is a section marked BEGIN Cloudflare Managed content, listing user agents like Amazonbot, Bytespider, CCBot, ClaudeBot, Google-Extended and GPTBot, each with Disallow: /.

How to check what crawlers really receive

Do not test from your own browser cache. Fetch the file the way a crawler would, from outside your network.

The quickest reliable check is a tool that fetches from its own servers — our free AI visibility checker reports each AI crawler's access individually and quotes the exact rules blocking it. You can also ask someone on a different network to open yoursite.com/robots.txt, or use your phone on mobile data instead of Wi-Fi.

Look specifically for the Cloudflare managed block. If it is there, your own rules further down the file will not rescue you, because a crawler obeys the first matching group for its name.

Turning it off in Cloudflare

Log in to the Cloudflare dashboard and select your domain. Open AI Crawl Control, which sits under the Security or Bots area depending on your plan.

Set AI crawlers to Allow, and switch off the managed robots.txt option so Cloudflare stops injecting its rules into your file.

Changes take effect within minutes. Re-run a scan afterwards to confirm the served robots.txt no longer contains the managed block.

One caveat worth understanding: this feature exists because some site owners genuinely want to keep AI models away from their content. If your business sells content, blocking may be the right call. If your business needs to be recommended, it is not.

Frequently asked questions

Does Cloudflare block AI bots by default?

▾

It depends on your plan and when your zone was created. Cloudflare has progressively made blocking AI crawlers the default for new domains, so many site owners have it switched on without ever having chosen it. Always verify rather than assume.

Will allowing AI crawlers slow my site or raise costs?

▾

In practice the traffic is negligible for most sites — AI crawlers fetch far less aggressively than search engine crawlers. If you serve very large media files, you can allow the crawlers for your HTML pages while disallowing heavy asset directories.

Do other CDNs do this too?

▾

Cloudflare is the most prominent, but similar bot-management features exist at Fastly, Akamai and several WordPress hosts. If your robots.txt looks correct at the origin but crawlers still report a block, check whatever sits in front of your server.