Blog · · By HeardOf
Is your CDN blocking AI crawlers? What Cloudflare's defaults do, per its own pages
The short answer
It can. Cloudflare's Block AI Bots page, last updated 1 July 2026 and read 24 September 2026, says domains added from 15 September 2026 will start with Training and Agent blocked on pages with ads, Search allowed.
Your robots.txt is a request. What answers a crawler can be the CDN in front of your origin, and on Cloudflare that answer is a setting with a default. We read Cloudflare's own pages on 24 September 2026 — a press release and a blog post of 1 July 2025 and seventeen documentation pages, each dated below — and quote what they say about those defaults. We sent no crawler to any Cloudflare zone; the one measurement here is what a plain fetch with no JavaScript got from three of 30 B2B pricing pages the same day, and it says nothing about what GPTBot got.
Is Cloudflare blocking GPTBot on your site by default?
For a domain added from 15 September 2026, Cloudflare's documentation says Training bots will be blocked on pages that display ads, and GPTBot is a training crawler by our reading of two of its pages; for an older domain, the answer is whatever was chosen at sign-up or since. The page titled Block AI Bots, last updated 1 July 2026 and read 24 September 2026, states: "On September 15, 2026, Cloudflare will set updated defaults for new domains: bots classified as Training or as Agent will be blocked on pages that display ads, and Search will remain allowed." The bots changelog, in an entry dated 1 July 2026, says those defaults "take effect for new domains on September 15, 2026". Neither page had been put in the past tense when we read them nine days after the date, and no changelog entry is dated on or after it. So whether Cloudflare is blocking GPTBot on your site is a question its pages answer, as read, only for a domain created after mid-September, and even there in the future tense.
Where GPTBot falls is our inference from two other pages; Cloudflare prints no such mapping. The bot reference for AI Crawl Control, last updated 23 April 2026 and historical by our 90-day rule, lists 20 crawlers from 12 operators and puts GPTBot under the legacy category AI Crawler, which the verified-bots page defines as "Crawls websites for content that is used for training AI models", with "ChatGPT bot" as its example; OAI-SearchBot sits under AI Search, which that page says is, since 1 July 2026, "treated as Search behavior". Our reading of the two together, not a sentence on either page, puts OpenAI's training crawler in the set a new domain would block on ad pages and its search crawler in the set it would allow. Neither page prints a per-crawler mapping to the three presets, and the verified-bots page says a bot can have "one or more" behaviours, classified in BotBase, which we did not read.
The word default is older than the presets. The press release of 1 July 2025, historical by our rule, says Cloudflare "is now the first Internet infrastructure provider to block AI crawlers accessing content without permission or compensation, by default" and that "every new domain will now be asked if they want to allow AI crawlers"; the blog post of the same day calls it "changing the default to block AI crawlers unless they pay creators for their content".
What does the Block AI bots setting block, and what replaced it?
The older toggle "blocks verified bots that are classified as crawling for the purpose of AI training, as well as a number of unverified bots that behave similarly", "excludes mixed-purpose bots that are used both for Training and for Search", and is marked "Deprecating on September 15, 2026". Its replacement is one preset per behaviour, each with three options. The behaviours, from the bots concepts page of 1 July 2026: Search "Collects or indexes your content so it can answer questions about it later"; Agent, "Automated activity acting in real time on a person's behalf to get something done, such as chat fetch bots and browser-use agents"; Training, "Crawls your content to train or fine-tune a model, permanently absorbing your data into the model". Each preset "will block Verified bots classified with that behavior, plus additional unverified bots that fall under these classifications", and lives under Security Settings at the entry labelled Configure AI bot policies. From 15 September 2026, mixed-purpose crawlers "will also be blocked by all configurations to block AI training, including the legacy "Block AI bots" option".
Does Bot Fight Mode block GPTBot or PerplexityBot?
Not by name on the pages we read, and the verified-bots page says verified bots were historically excluded from default bot configurations. Bot Fight Mode, "a simple, free product", "Identifies traffic matching patterns of known bots" and "Issues computationally expensive challenges that force the requesting client to perform CPU-intensive calculations"; for its users "JavaScript Detections is automatically enabled and cannot be disabled". Neither its page nor Super Bot Fight Mode's names GPTBot, OAI-SearchBot or PerplexityBot; each sends the reader to Block AI bots. The verified-bots page: "Historically, Verified bots have been excluded in default bot configurations across all plans. Now, all customers have the option to configure AI bot policies to define their block vs. allow expectations." Which of the 20 crawlers in the bot reference is Verified is not printed there; it points to Cloudflare Radar's bots directory "For an up-to-date list of verified bots", which we did not read.
What does a client that is not a browser actually get?
A challenge page or a 403, from 3 of the 30 B2B pricing pages we read on 24 September 2026, and none of the three tells you what an AI crawler got. Vincere's pricing page answered a plain fetch — a Chrome 140 user-agent string, no JavaScript — with HTTP 403 and a Cloudflare challenge page reading "Just a moment... Enable JavaScript and cookies to continue", a response header naming the mitigation as a challenge, at 17:43:32 UTC; headless Chrome was pointed at the same URL once, and what it received we do not report. Fleetio's answered both clients with a 403 from DataDome, and Patterson Dental's with a plain 403 from an Azure gateway. The other 27 are in 30 B2B software pricing pages and JavaScript. The same day, 2 of the 88 domains ChatGPT had cited in our run of 4 September 2026, and 7 of the 204 Perplexity had cited, answered a request for robots.txt with a Cloudflare challenge page, counted in do the sites ChatGPT and Perplexity cite block their crawlers?.
None of this says an AI crawler was blocked at any of those sites. We sent none; a Chrome user-agent string with no script is not GPTBot, and the verified-bots page says a verified bot identifies itself "through a cryptographic Web Bot Auth signature, a published IP list with a stable user-agent, or reverse DNS". OpenAI's and Perplexity's crawler pages, read the same day for AI crawlers in one table, do not say what their bots do when they meet a challenge. How much those crawlers take relative to what they send back is a separate count, and the sources that publish it do not share a unit: Cloudflare's crawl-to-refer ratios, TollBit's scrape-to-referral figures, Fastly's volume shares and Vercel's request counts are compiled, each with its unit, window and date, in AI crawlers take more than they send back.
How do you check what your Cloudflare zone does to AI crawlers?
In two areas of the dashboard, per the docs: Security Settings, which holds the AI bot policies, Bot fight mode and the managed robots.txt, and AI Crawl Control, which holds the per-crawler actions. AI Crawl Control is "Available on all plans"; its Crawlers tab lists each crawler with an Action column offering Allow or Block, and "When you block a crawler in AI Crawl Control, the system creates or updates a WAF custom rule on your zone to enforce that block"; on the free plan it "identifies AI crawlers based on their user agent strings". Perplexity's crawler page, read 24 September 2026, assumes that layering — "you may need to explicitly whitelist Perplexity's bots" — and its Cloudflare steps are a custom rule under Security, then WAF, matching the user agent and its published IP ranges, with the action set to Allow. The managed robots.txt disallows eight named crawlers on every path, GPTBot among them, and its page says "robots.txt compliance is voluntary"; the eight are named in the questions below.
| Setting, as the docs label it | Where the docs say it lives | What the page says it does | Plans the page names | Date the page prints |
|---|---|---|---|---|
| Configure AI bot policies | Security Settings, entry Configure AI bot policies | three presets, Search, Agent and Training, each set to "Block (on all pages)", "Block on pages with ads" or "Allow (do not block)"; new-domain default Cloudflare says it will set from 15 September 2026: Training and Agent blocked on pages with ads, Search allowed | "All Cloudflare customers" | 1 July 2026 |
| Block AI bots | Security Settings, entry Block AI bots | "blocks verified bots that are classified as crawling for the purpose of AI training, as well as a number of unverified bots that behave similarly"; excludes mixed-purpose bots; "Deprecating on September 15, 2026" | none stated | 1 July 2026 |
| Bot fight mode | Security Settings, filtered by Bot traffic | "Identifies traffic matching patterns of known bots" and "Issues computationally expensive challenges"; JavaScript Detections on and "cannot be disabled"; cannot be skipped by WAF custom rules | "a simple, free product" | 3 August 2026 |
| Super Bot fight mode | Security Settings, filtered by Bot traffic | per-grouping values for Definitely automated, Likely automated and Verified bots; "Can challenge or block bots"; exceptions by a custom rule with the Skip action | Pro, Business, Enterprise | 3 August 2026 |
| AI Crawl Control | AI Crawl Control, Crawlers tab, Action column | per-crawler Allow or Block, enforced by one WAF custom rule; free plan detects by user-agent string; block response 403 or 402 configurable on paid plans | "Available on all plans" | 14 August 2026 (overview); 28 July 2026 (manage) |
| Set your preference to block training in robots.txt | Security Settings, filtered by Bot traffic | prepends a managed robots.txt that disallows 8 named crawlers on every path and carries a content-signal line search yes, ai-train no, use reference; "robots.txt compliance is voluntary" | "available on all plans" | 3 August 2026 |
| AI Labyrinth | Security Settings, filtered by Bot traffic | "adds invisible links on your webpage with specific Nofollow tags"; "AI Labyrinth actions are not mitigations. Cloudflare does not block or challenge the request." | none stated | 17 September 2026 |
| Pay per crawl | AI Crawl Control, Action column, Charge | HTTP 402 with pricing unless the crawler presents payment intent; a WAF or Bot Management block overrides the charge | "closed beta"; no plan named beyond sending existing Enterprise customers to their account executive (historical Get started page, 23 April 2026: charging listed under Enterprise plans with Bot Management) | 28 July 2026 |
How was this read?
With one plain fetch per page on 24 September 2026 between 22:23 and 22:33 UTC, a Chrome 140 user-agent string and no JavaScript, from one machine in one location, then a render in headless Chrome 153 of fourteen of them between 22:26 and 22:40 UTC, and of the other five — AI Crawl Control with Cloudflare Bots, What is Pay Per Crawl?, the AI Crawl Control changelog, AI Labyrinth and the Bots overview — between 22:58 and 23:00 UTC. Every page answered 200 both ways, none showed a bot challenge, nothing was bypassed, and the dates are what each page prints. Two are 450 days old and historical by our 90-day rule: the press release and the blog post, both 1 July 2025. The documentation: Block AI Bots, Bots, Verified bots and AI Crawl Control with Cloudflare Bots, 1 July 2026; Manage AI crawlers and What is Pay Per Crawl?, 28 July; Bot Fight Mode, Super Bot Fight Mode, Managed robots.txt and AI Crawl Control with Cloudflare WAF, 3 August; the AI Crawl Control overview, 14 August; AI Labyrinth, 17 September; and, historical, Get started and the bot reference, 23 April, the Bots overview, 30 April, and the AI Crawl Control changelog and Bots changelog, April dates with newest entries of 16 June and 1 July 2026. Perplexity's crawler page prints no date. Each documentation page opens its served HTML with a line addressed to agents, "Use this file to discover all available pages before exploring further", which we read as data and did not follow; nor did we follow the opt-out link on the Block AI Bots page, which points at the Cloudflare dashboard.
The measurement here is ours: what three pricing pages and robots.txt requests to 269 domains returned to our client on 24 September 2026, from our own fetch logs of that day, which you cannot re-run as they were. Everything else is a quotation from a Cloudflare or Perplexity page, linked above with the date it prints and the day we read it, and a page can change the day after it is read. We sent no AI crawler, and nothing here says what one received.
Common questions
Does Cloudflare block AI crawlers by default?
For a domain added from 15 September 2026, its Block AI Bots page, last updated 1 July 2026 and read 24 September 2026, says bots classified as Training or Agent "will be blocked on pages that display ads, and Search will remain allowed"; the page still says "will" on the day we read it, and neither changelog we read has an entry on or after that date. For a domain added between 1 July 2025 and then, the press release of 1 July 2025, historical by our 90-day rule, says every new domain "will now be asked if they want to allow AI crawlers" at sign-up. For an older domain, the pages we read name no new default; the verified-bots page says "Historically, Verified bots have been excluded in default bot configurations across all plans", and the Block AI Bots page says mixed-purpose crawlers will also be blocked by any configuration that already blocks AI training.
What is the difference between Block AI bots and AI Crawl Control?
Block AI bots is a single toggle under Security Settings that "blocks verified bots that are classified as crawling for the purpose of AI training", marked "Deprecating on September 15, 2026" on the Block AI Bots page of 1 July 2026, read 24 September 2026; since 1 July 2026 that page also offers one preset per behaviour — Search, Agent, Training — under Configure AI bot policies. AI Crawl Control is per crawler: an Action column offering Allow or Block, and "When you block a crawler in AI Crawl Control, the system creates or updates a WAF custom rule on your zone to enforce that block". The two run at different points: the docs say WAF custom rules "take place before Cloudflare bot solutions", and that to make pay per crawl work "you need to first turn off Block AI Bots".
Does Cloudflare's managed robots.txt block GPTBot?
It asks it to stay out. The page, last updated 3 August 2026 and read 24 September 2026, prints the managed file: a Disallow line for the root path under each of eight named crawlers — Amazonbot, Applebot-Extended, Bytespider, CCBot, ClaudeBot, Google-Extended, GPTBot and meta-externalagent — and says "robots.txt compliance is voluntary" and "If you want to enforce crawl blocking rather than request it, use AI Crawl Control". OAI-SearchBot, ChatGPT-User, PerplexityBot and Perplexity-User are not in the managed list.
Why would a crawler I allowed still get a 403 from my Cloudflare site?
Cloudflare's page on AI Crawl Control with Cloudflare WAF, last updated 3 August 2026 and read 24 September 2026: "check for upstream WAF custom rules that may be blocking them", because the AI Crawl Control rule "only includes blocked bots" and "is added at the end of existing WAF custom rules". Bot Fight Mode is a separate case, per its page of 3 August 2026: "You cannot bypass or skip Bot Fight Mode using WAF custom rules or Page Rules", and its page says that for Bot Fight Mode customers "JavaScript Detections is automatically enabled and cannot be disabled". Which code a block returns before anyone edits it — the docs offer 403 Forbidden and 402 Payment Required on paid plans — is not stated on the pages we read.
What is pay per crawl?
A feature of AI Crawl Control "currently in closed beta", per its page of 28 July 2026 read 24 September 2026: a crawler either presents "payment intent via request headers for successful HTTP 200 access", or receives "an HTTP 402 Payment Required response with pricing", with Cloudflare "as the Merchant of Record". A crawler blocked by a WAF or Bot Management rule is not charged; "the blocked crawler will not have access to the zone". It is one of three actions in the Action column — Allow, Charge, Block — on accounts in the beta.
These are our numbers. Yours are one audit away.
Nineteen Cloudflare pages and one of Perplexity's, read on one day and quoted with the date each prints, because a default is a sentence a vendor can rewrite. Whether a crawler reaches your pages is a question for your CDN's dashboard and your logs. Whether the engines those crawlers feed name you is a different one — 40 buyer prompts, you against up to three competitors that you name, as read on 19 September 2026 in our post on which engines we check — and that is what HeardOf does.
AI crawlers take more than they send back, on the latest figures: crawl-to-referral ratios, compiled