Blog · · By HeardOf
Do the sites ChatGPT and Perplexity cite block their crawlers?
The short answer
Rarely. Of the 269 domains ChatGPT and Perplexity cited in our run of 20 prompts on 4 September 2026, 253 served a readable robots.txt on 24 September 2026, and 7 disallow an OpenAI or Perplexity bot on every path.
Twenty days after the run, we went back to the sources. On 4 September 2026 we asked ChatGPT and Perplexity 20 buyer prompts in four B2B software categories and kept every URL they returned as a source: 564 citations across 269 domains, the run counted in who the AI actually cites. On 24 September 2026 we requested each of those domains' robots.txt and read what it says about five of the bots OpenAI and Perplexity document — GPTBot, OAI-SearchBot, ChatGPT-User, PerplexityBot and Perplexity-User; OpenAI's page also lists OAI-AdsBot, which it says visits only pages submitted as ads, and we did not count it. The two dates stay apart throughout: nothing here says what an engine did on 24 September or what a file said on 4 September.
Do the sites ChatGPT cites block GPTBot?
Three of them do, out of the 81 whose robots.txt we could read on 24 September 2026, and none of the 81 disallows OAI-SearchBot on every path. The three — gpstrackit.com, news.remax.com and optimoroute.com — were each cited once by ChatGPT on 4 September 2026, and each carries a group naming GPTBot followed by a Disallow line for the root path, read 24 September 2026.
Each leaves the search crawler open: gpstrackit.com names OAI-SearchBot with an Allow line for the root followed by five preview, preference and tracking paths, and the other two do not name it, so it falls under their rule for all crawlers, which disallows one topic path at news.remax.com and nothing at optimoroute.com. news.remax.com also disallows ChatGPT-User on every path, and optimoroute.com also disallows PerplexityBot. Nine of the 81 name any of the five bots at all: eight name GPTBot, five OAI-SearchBot. Seven of ChatGPT's 88 domains are outside the count: business.adobe.com did not answer on three attempts; inman.com and thedigitalprojectmanager.com answered a Cloudflare challenge page, which counts as blocked to us and was not bypassed; s3.amazonaws.com, sec.gov and sierrainteractive.com answered 403; support.teamwork.com answered 404.
Does blocking GPTBot stop ChatGPT citing you?
Not according to the document OpenAI publishes, and not in what we can see, which is less than the question asks. The page 'Overview of OpenAI Crawlers', read 24 September 2026, gives the two bots different jobs: GPTBot "is used to crawl content that may be used in training our generative AI foundation models", and disallowing it "indicates a site's content should not be used in training"; OAI-SearchBot "is used to surface websites in search results in ChatGPT's search features", and "sites that are opted out of OAI-SearchBot will not be shown in ChatGPT search answers, though can still appear as navigational links". Each setting is independent, a robots.txt change takes about 24 hours to reach its systems, and for ChatGPT-User, the agent that may fetch a page when a user asks, "robots.txt rules may not apply".
That sentence is about OAI-SearchBot, and the page spells the pairing out: a webmaster "can allow OAI-SearchBot in order to appear in search results while disallowing GPTBot". So the answer to does blocking GPTBot stop ChatGPT citing me, on OpenAI's page, is no. Our data agrees as far as it goes and no further: three sites that disallowed GPTBot on 24 September were cited on 4 September, we did not ask ChatGPT again on 24 September, we do not know what those files said on 4 September, and three domains at one citation each is a photograph, not a rate. What the data does not show is the reverse: no site ChatGPT cited that we could read had, twenty days later, a rule disallowing, on every path, the bot OpenAI says governs search.
Do the sites Perplexity cites block PerplexityBot?
Three of the 193 we could read, and 19 of the 21 citations involved are reddit.com. techstackdaily.com names PerplexityBot, with GPTBot, ChatGPT-User and eighteen other user agents, in one group above a Disallow line for the root path, under a comment dated 8 September 2026 — four days after Perplexity cited the site once, on 4 September 2026. reddit.com and support.billsby.com name none of the five and disallow every path for all user agents, so every bot they do not name is disallowed; Perplexity cited reddit.com 19 times and support.billsby.com once in the same run. Reddit's file, read 24 September 2026, is two lines under five comment lines, two of them pointing at Reddit's Public Content Policy: a rule for all user agents and a Disallow line for the root path.
Perplexity's crawler page, read the same day, says PerplexityBot "is designed to surface and link websites in search results on Perplexity" and recommends allowing it, and that Perplexity-User, which "might visit a web page to help provide an accurate answer and include a link to the page in its response", "generally ignores robots.txt rules". A citation does not say which agent fetched the page, so we cannot tell you which one produced the 19, and whether Reddit and Perplexity have an agreement that makes the file beside the point is a question we did not check. The number stands as what it is: on 4 September 2026, 21 of Perplexity's 391 citations, 5.4%, pointed at a domain whose robots.txt on 24 September 2026 disallowed PerplexityBot on every path. Eleven of Perplexity's 204 domains could not be read — seven served a Cloudflare challenge page, linkedin.com a reCAPTCHA page in place of a robots file, three a 403 — and are named in our research file.
How many cited sites name an OpenAI or Perplexity bot at all?
68 of the 253 we could read, 26.9%, name at least one of the five bots, and what the named rules mostly say is allow. GPTBot is named in 62 files, PerplexityBot in 60, ChatGPT-User in 49, OAI-SearchBot in 39 and Perplexity-User in 19; 16 files name all five. 28 files name GPTBot and say nothing about OAI-SearchBot, the bot OpenAI's page says governs search; 41 name PerplexityBot and nothing about Perplexity-User.
| Bot | Named, of 253 files | Named, allows every path | Named, disallows some paths | Named, disallows every path | Disallowed on every path in effect, named or by a rule for all crawlers | Of ChatGPT's 81 domains, disallowed in effect | Of Perplexity's 193 domains, disallowed in effect |
|---|---|---|---|---|---|---|---|
| GPTBot | 62 | 34 | 24 | 4 | 7 | 3 | 4 |
| OAI-SearchBot | 39 | 19 | 20 | 0 | 2 | 0 | 2 |
| ChatGPT-User | 49 | 29 | 18 | 2 | 4 | 1 | 3 |
| PerplexityBot | 60 | 35 | 23 | 2 | 4 | 1 | 3 |
| Perplexity-User | 19 | 6 | 13 | 0 | 2 | 0 | 2 |
The two engines' pools differ here too: nine of the 81 ChatGPT-cited files name any of the five bots, 11.1%, against 63 of the 193 Perplexity-cited files, 32.6%. The pools are different kinds of site — in the same run, 44.5% of ChatGPT's citations pointed at a measured vendor's own site against 9.2% of Perplexity's, as counted in the post linked above — and a vendor's own file, on this read, rarely names the bots. Of the 30 B2B software brands in our 4 September 2026 measurement (named but not cited) — six of them dental brands whose stored answers were later lost — 26 served a readable robots.txt on 24 September 2026 and three name any of the five. Of the 19 AI-visibility vendors in our pricing post, 18 did and two name any of the five; none of those five files disallows a bot on every path. The brand and vendor files we could not read, and the two cited files that mention the bots only inside comments, are in the research file.
Are the cited sites in Tranco's top 100,000?
Most of Perplexity's are not, and most of ChatGPT's are. Against Tranco list L5PV4 — generated 23 September 2026 from Crux, Farsight, Majestic, Radar and Umbrella over the 30 days to that date, downloaded 24 September 2026 — 34 of the 88 domains ChatGPT cited on 4 September 2026 rank outside the top 100,000, 38.6%, against 149 of Perplexity's 204, 73.0%. Per citation rather than per domain: 64 of ChatGPT's 173, 37.0%, and 240 of Perplexity's 391, 61.4%. Not in its top 1,000,000: 4 of ChatGPT's 88 domains and 84 of Perplexity's 204. Across both engines, 173 of the 269 domains, 64.3%, are outside the top 100,000.
The Perplexity figure extends one that exists. Trellner's study, read 24 September 2026, put 380 software categories to perplexity/sonar and perplexity/sonar-pro through OpenRouter on 2 September 2026 and kept 7,534 citations: 59.8% of them point at a domain ranked worse than 100,000 in Tranco's daily list of 1 September 2026, and 23.4% at a domain not in its top million; only Perplexity was measured, one prompt wording, one run per category. Its unit is the citation, so the like-for-like pair is 61.4% of Perplexity's 391 citations in our run against its 59.8% — two days apart, on different prompts, against different Tranco lists. Our 73.0% of domains outside the top 100,000 has no counterpart on its page; its 36.5% of domains not in the top million sits beside our 41.2% of Perplexity's 204 domains. What one engine could not give is ChatGPT on the same prompts the same day: 37.0% of its citations and 38.6% of its domains outside the top 100,000. Tranco ranks pay-level domains, so a subdomain carries its parent's rank; reddit.com ranks 109.
How was this measured?
The 269 domains are the hosts of the 564 citations stored from the run of 4 September 2026, with www. removed at the start and subdomains kept, the rule of the post on who the AI cites; re-derived on 24 September 2026, the count matches that post's 269, 88 for ChatGPT and 204 for Perplexity with 23 shared. On 24 September 2026, between 18:40 and 18:45 UTC, each domain's robots.txt was requested over https with a Chrome user agent and no JavaScript, from one machine in one location, redirects followed; twenty domains were re-read with a second client, and two whose certificate does not match their name were read over plain http, as our fetch log records. A 403, a challenge page or a 200 whose body is a reCAPTCHA page is recorded as blocked to us and was not retried with a browser; 15 of the 269 ended that way or did not answer, and one, support.teamwork.com, has no file. Each file was read by the grouping rule of the robots exclusion standard — a bot's rules are the lines in the group that names it, or the group for all user agents when none does — and a group disallows every path when a Disallow line covers the root and no Allow line in the group reopens any path, which is the case in all seven. Perplexity's sentences were taken from the markdown copy the site serves, and match its page as served. Every per-domain row is in our research file, and every total here was computed from those rows.
Everything this post says about our own run — which domains were cited, how many times, by which engine — comes from our own stored answers of 4 September 2026 as re-read on 24 September 2026, and carries that date in the sentence that makes it; that is the half of this post you cannot check, and we are saying so rather than letting you find the seam. Everything else carries a link or a date you can follow: OpenAI's crawler page and Perplexity's, both read 24 September 2026; Tranco list L5PV4, generated 23 September 2026; Trellner's study, dated 2 September 2026; and every robots.txt named here, which you can request yourself today and which may no longer say what it said on 24 September 2026.
Common questions
Does blocking GPTBot stop ChatGPT from citing my site?
Not according to OpenAI's own crawler page, read 24 September 2026, which describes GPTBot as the crawler for training and says that a site opted out of OAI-SearchBot "will not be shown in ChatGPT search answers". Our data points the same way with a caveat: three of the 81 ChatGPT-cited domains whose robots.txt we could read disallow GPTBot on every path, and all three were cited on 4 September 2026, twenty days before we read the files. We did not ask ChatGPT again on 24 September, and a robots.txt can change in twenty days.
What is the difference between GPTBot and OAI-SearchBot?
OpenAI's page, read 24 September 2026, says GPTBot "is used to crawl content that may be used in training our generative AI foundation models" and that disallowing it "indicates a site's content should not be used in training". OAI-SearchBot "is used to surface websites in search results in ChatGPT's search features", and OpenAI recommends allowing it. The page says each setting is independent and that a robots.txt change takes about 24 hours to reach its systems. Of the 253 cited-domain robots files we could read, 62 name GPTBot and 39 name OAI-SearchBot; 28 name GPTBot and say nothing about OAI-SearchBot.
Does Perplexity respect robots.txt?
Its crawler page, read 24 September 2026, gives two answers for two agents. PerplexityBot "is designed to surface and link websites in search results on Perplexity", and Perplexity recommends allowing it. Perplexity-User may fetch a page when a user asks a question and, the page says, "generally ignores robots.txt rules". In our run of 4 September 2026, 21 of Perplexity's 391 citations went to a domain whose robots.txt on 24 September 2026 disallows PerplexityBot on every path, 19 of them reddit.com; which agent fetched those pages is not something a citation shows.
Does Reddit's robots.txt block AI crawlers?
It disallows every crawler it does not name, and it names none: read 24 September 2026, the file is a rule for all user agents and a Disallow line for the root path, under five comment lines, two of them pointing at Reddit's Public Content Policy. It does not mention GPTBot, OAI-SearchBot, ChatGPT-User, PerplexityBot or Perplexity-User. Perplexity cited reddit.com 19 times on 4 September 2026 and ChatGPT zero times. Whether Reddit has an agreement with either company is not something we checked.
These are our numbers. Yours are one audit away.
Two dates, kept apart: what the robots.txt files of 253 of the 269 cited domains said on 24 September 2026, and what two engines cited on 4 September 2026. If what you want measured is not what your robots.txt says but whether the engines name you — 40 buyer prompts, you against up to three competitors that you name, as read on 19 September 2026 in our post on which engines we check — that is what HeardOf does.