Blog · · By HeardOf
AI crawlers in one table: what OpenAI, Anthropic, Perplexity, Google, Apple, Common Crawl and Meta document
The short answer
Twenty-nine crawler and fetcher names across seven vendor pages, read 24 September 2026. Five carry a stated exception to robots.txt, three of them the agents that fetch a page when a user asks. One page mentions JavaScript.
Seven companies publish a page about their own web crawlers. We read all seven on 24 September 2026 and put what they say side by side: the user-agent name each bot uses, what the vendor says it is for, and what the vendor says about robots.txt, quoted rather than paraphrased. The pages are the source. We did not watch the bots, and nothing here says what a crawler does, only what its owner writes.
Which AI crawler user agents does each vendor document?
Twenty-nine names across the seven pages, read 24 September 2026: four from OpenAI, two from Perplexity, three from Anthropic, eleven from Google, three from Apple, one from Common Crawl and five from Meta. The table carries 22 of them; Google's other seven — Googlebot-Image, Googlebot-Video, Googlebot-News, Storebot-Google, Google-InspectionTool, GoogleOther-Image and GoogleOther-Video — are for images, video, news, shopping, Search Console's testing tools and the image and video forms of GoogleOther, and Google's page says its list "is not exhaustive". Anthropic's page names three bots and prints a user-agent string for none of them. The other six print one for every agent that fetches pages; Google-Extended, Googlebot-News and Applebot-Extended are names without a string of their own.
| Vendor | Agent | User-agent string as printed on the page | What the vendor says it is for | What the vendor says about robots.txt | Date the page shows |
|---|---|---|---|---|---|
| OpenAI | OAI-SearchBot | Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/131.0.0.0 Safari/537.36; compatible; OAI-SearchBot/1.4; +https://openai.com/searchbot — an example, "the version number may change" | "used to surface websites in search results in ChatGPT's search features" | "we recommend allowing OAI-SearchBot in your site's robots.txt file"; opted-out sites "will not be shown in ChatGPT search answers" | none |
| OpenAI | OAI-AdsBot | Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; OAI-AdsBot/1.0; +https://openai.com/adsbot | "used to validate the safety of web pages submitted as ads on ChatGPT"; "only visits pages submitted as ads" | the page says nothing about robots.txt for this agent | none |
| OpenAI | GPTBot | Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.4; +https://openai.com/gptbot — an example | "used to crawl content that may be used in training our generative AI foundation models" | "Disallowing GPTBot indicates a site's content should not be used in training generative AI foundation models" | none |
| OpenAI | ChatGPT-User | Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; ChatGPT-User/1.0; +https://openai.com/bot | "may visit a web page with a ChatGPT-User agent" when a user asks; "not used for crawling the web in an automatic fashion" | "Because these actions are initiated by a user, robots.txt rules may not apply" | none |
| Perplexity | PerplexityBot | Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; PerplexityBot/1.0; +https://perplexity.ai/perplexitybot) | "designed to surface and link websites in search results on Perplexity"; "not used to crawl content for AI foundation models" | "we recommend allowing PerplexityBot in your site's robots.txt file" | none |
| Perplexity | Perplexity-User | Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Perplexity-User/1.0; +https://perplexity.ai/perplexity-user) | "might visit a web page to help provide an accurate answer and include a link to the page in its response" | "Since a user requested the fetch, this fetcher generally ignores robots.txt rules" | none |
| Anthropic | ClaudeBot | none printed | "collecting web content that could potentially contribute to their training" | restricting it "signals that the site's future materials should be excluded from our AI model training datasets"; the page says its bots respect do-not-crawl signals "by honoring industry standard directives in robots.txt" | April 7, 2026, printed without a label; 170 days before our read, historical by our 90-day rule |
| Anthropic | Claude-User | none printed | "When individuals ask questions to Claude, it may access websites using a Claude-User agent" | "Disabling Claude-User on your site prevents our system from retrieving your content in response to a user query" | April 7, 2026 |
| Anthropic | Claude-SearchBot | none printed | "navigates the web to improve search result quality for users" | disabling it "prevents our system from indexing your content for search optimization" | April 7, 2026 |
| Googlebot | Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Googlebot/2.1; +http://www.google.com/bot.html) Chrome/W.X.Y.Z Safari/537.36 — the desktop form; W.X.Y.Z stands for the Chrome version | preferences "affect Google Search (including Discover and all Google Search features)" and Google Images, Video and News | the common crawlers "always obey robots.txt rules when crawling automatically" | Last updated 2026-07-14 UTC | |
| Google-Extended | none: "doesn't have a separate HTTP request user agent string"; the robots.txt name "is used in a control capacity" | controls whether crawled content "may be used for training future generations of Gemini models" and "for grounding" in Gemini Apps and Vertex AI | "does not impact a site's inclusion in Google Search nor is it used as a ranking signal in Google Search" | 2026-07-14 | |
| GoogleOther | Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; GoogleOther) Chrome/W.X.Y.Z Safari/537.36 | "the generic crawler that may be used by various product teams", for example "one-off crawls for internal research and development" | as Googlebot, by the page-level sentence | 2026-07-14 | |
| Google-CloudVertexBot | the substring Google-CloudVertexBot | "crawls requested by the site owners' for building Vertex AI Agents"; "no effect on Google Search or other products" | as Googlebot, by the page-level sentence | 2026-07-14 | |
| Apple | Applebot | Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/605.1.15 (KHTML, like Gecko) Version/17.4 Safari/605.1.15 (Applebot/0.1; +http://www.apple.com/go/applebot) — the desktop example | powers search in "Spotlight, Siri, and Safari"; its data "may also be used to help train Apple foundation models" | "respects standard robots.txt directives in general search crawls that are targeted at Applebot"; follows Googlebot rules when not named; "does not follow crawl-delay" | Published Date: September 04, 2026 |
| Apple | Applebot-Extended | none: "does not crawl webpages" | lets publishers "opt out of their website content being used to train Apple's general purpose foundation models" | pages that disallow it "can still be included in search results" | September 04, 2026 |
| Apple | iTMS | iTMS | "only crawls URLs associated with registered content on Apple Podcasts" | "does not follow robots.txt, as it is not a general search crawler" | September 04, 2026 |
| Common Crawl | CCBot | CCBot/2.0 (https://commoncrawl.org/faq/) | maintains "an open repository of web crawl data that is universally accessible and analyzable by anyone" | to prevent crawling, a group naming CCBot with a Disallow line for the root path; "we are aware of crawlers falsely identifying themselves as CCBot" | none |
| Meta | FacebookExternalHit | facebookexternalhit/1.1 (+http://www.facebook.com/externalhit_uatext.php) | "crawl the content of an app or website that was shared on one of Meta's family of apps" | "might bypass robots.txt when performing security or integrity checks" | Updated: May 21, 2026; 126 days before our read, historical by our 90-day rule |
| Meta | Meta-WebIndexer | meta-webindexer/1.1 (+/documentation/sharing/webmasters/web-crawlers) | "navigates the web to improve Meta AI search result quality for users" | allowing it "helps us cite and link to your content in Meta AI's responses" | May 21, 2026 |
| Meta | Meta-ExternalAds | meta-externalads/1.1 (+/documentation/sharing/webmasters/web-crawlers) | "improving advertising and other business-related products and services" | the page's rule for all five: "add a disallow for the relevant crawler to robots.txt" | May 21, 2026 |
| Meta | Meta-ExternalAgent | meta-externalagent/1.1 (+/documentation/sharing/webmasters/web-crawlers) | "training foundation AI models or improving products by indexing content directly" | the page's rule for all five, and its example group names this agent | May 21, 2026 |
| Meta | Meta-ExternalFetcher | meta-externalfetcher/1.1 (+/documentation/sharing/webmasters/web-crawlers) | "fetches individual links at a user's request" | "this crawler may bypass robots.txt rules" | May 21, 2026 |
What is the difference between GPTBot and OAI-SearchBot?
GPTBot is OpenAI's training crawler and OAI-SearchBot is its search crawler, and OpenAI's page says each robots.txt setting "is independent of the others". GPTBot "is used to crawl content that may be used in training our generative AI foundation models", and disallowing it "indicates a site's content should not be used in training generative AI foundation models". OAI-SearchBot "is used to surface websites in search results in ChatGPT's search features", and "Sites that are opted out of OAI-SearchBot will not be shown in ChatGPT search answers, though can still appear as navigational links".
So GPTBot vs OAI-SearchBot is not a choice between two names for one bot. A site can disallow the first and allow the second, and the page gives that as its own example of what the two settings are for. Two more details from the same page: if a site allows both, OpenAI "may use the results from just one crawl for both use cases to avoid duplicative crawling", and a robots.txt change takes "~24 hours" to reach its search systems. The third OpenAI agent, ChatGPT-User, is neither of these. It "may visit a web page" when a user asks a question, "is not used for crawling the web in an automatic fashion", and "is not used to determine whether content may appear in Search"; the page sends anyone managing search back to OAI-SearchBot.
Which of these crawlers say they obey robots.txt, and which say they may not?
Five of the 29 names carry a stated exception on their own vendor's page, and three of those five are the agents that fetch a page because a user asked. ChatGPT-User: "Because these actions are initiated by a user, robots.txt rules may not apply." Perplexity-User: "Since a user requested the fetch, this fetcher generally ignores robots.txt rules." Meta-ExternalFetcher "fetches individual links at a user's request" and, the page continues, "Accordingly, this crawler may bypass robots.txt rules". The fourth user-requested agent on these seven pages, Claude-User, is the one whose vendor says the file works on it: "Disabling Claude-User on your site prevents our system from retrieving your content in response to a user query", under a page-level line that Anthropic's bots respect do-not-crawl signals "by honoring industry standard directives in robots.txt". The other two exceptions are not about users. FacebookExternalHit "might bypass robots.txt when performing security or integrity checks", and Apple's iTMS, which crawls Apple Podcasts content, "does not follow robots.txt, as it is not a general search crawler". Those counts cover the seven pages in the table. Google's common-crawlers page links two more lists we did not tabulate: its user-triggered fetchers, which include Google-Agent and, the page says, "generally ignore robots.txt rules" because "the fetch was requested by a user" (last updated 2026-08-19), and its special-case crawlers, which "may ignore robots.txt rules" (last updated 2026-09-17), both read 24 September 2026.
For the automatic crawlers in the table, the pages treat the file as the control, each in its own terms. Google: the common crawlers "always obey robots.txt rules when crawling automatically". Apple: Applebot "respects standard robots.txt directives in general search crawls that are targeted at Applebot", and "If robots instructions don't mention Applebot but mention Googlebot, the Apple robot will follow Googlebot instructions". Common Crawl gives the group to write, a rule naming CCBot with a Disallow line for the root path, and adds that it is "aware of crawlers falsely identifying themselves as CCBot". Meta: "add a disallow for the relevant crawler to robots.txt". Three of the seven pages say how long a change takes to land, and all three say a day: OpenAI "~24 hours", Perplexity "up to 24 hours", Meta "up to 24 hours". Google's page gives no figure, but the robots.txt page it links says Google "generally caches the contents of robots.txt file for up to 24 hours". Crawl-delay: Anthropic supports "the non-standard Crawl-delay extension to robots.txt", Apple's Applebot "does not follow crawl-delay", Common Crawl's FAQ says "We obey the Crawl-delay parameter for robots.txt", and Google's robots.txt page lists crawl-delay among the fields that "aren't supported". What the sites the engines cite actually write is a separate read: of the 269 domains ChatGPT and Perplexity cited in our own run of 4 September 2026, 253 served a readable robots.txt on 24 September 2026 and 7 disallow an OpenAI or Perplexity bot on every path, from our own stored answers, counted in do the sites ChatGPT and Perplexity cite block their crawlers?
Does a training opt-out also take you out of search or answers?
Not according to the five pages that separate the two controls; three say so outright and two describe separate bots. Google-Extended "does not impact a site's inclusion in Google Search nor is it used as a ranking signal in Google Search", and it "doesn't have a separate HTTP request user agent string": crawling "is done with existing Google user agent strings", and the robots.txt name "is used in a control capacity". Applebot-Extended "does not crawl webpages"; pages that disallow it "can still be included in search results", and it "is only used to determine how to use the data crawled by the Applebot user agent". OpenAI's, Anthropic's and Meta's training crawlers are separate bots that do crawl — GPTBot, ClaudeBot and Meta-ExternalAgent — each beside a search or answer crawler: OAI-SearchBot, Claude-SearchBot, and Meta-WebIndexer, which Meta says "helps us cite and link to your content in Meta AI's responses" when allowed.
The other two pages draw no such line. Perplexity's says PerplexityBot "is not used to crawl content for AI foundation models" and Perplexity-User "is not used for web crawling or to collect content for training AI foundation models". Common Crawl's says what the foundation maintains, "an open repository of web crawl data that is universally accessible and analyzable by anyone", and nothing about what the data trains. Apple adds a control that is not a user agent at all: the nosnippet meta tag, which its page says lets publishers "opt out of their content being used in these broad world knowledge answers" in Siri and Search.
Do any of these pages say whether the crawler runs JavaScript?
One of the seven, Apple's: "Applebot may render the content of your website within a browser. If javascript, CSS, and other resources are blocked via robots.txt, it may not be able to render the content properly." The words JavaScript, render and headless appear on none of the other six pages as read 24 September 2026, and the word browser appears on three of them only in navigation labels or in a note about a version placeholder. That includes OpenAI's and Perplexity's pages, which we first checked for 30 B2B software pricing pages and JavaScript earlier the same day and re-read for this post: neither says whether OAI-SearchBot or PerplexityBot executes a script, and neither says what its bot does when it meets a bot challenge. On the seven pages, only Apple's addresses rendering. Two vendors answer elsewhere: Common Crawl's FAQ, linked from its CCBot page, says "Currently, JavaScript is not executed", and Google's JavaScript SEO page, last updated 4 March 2026 and historical by our 90-day rule, says "a headless Chromium renders the page and executes the JavaScript". OpenAI, Perplexity, Anthropic and Meta say nothing about it on the pages we read. What answers a bot before the page does can be the CDN in front of it: Cloudflare's Block AI Bots page, last updated 1 July 2026 and read 24 September 2026, says that for domains added from 15 September 2026 it will block Training and Agent bots on pages that display ads by default and leave Search allowed, and what a client with no JavaScript got from three B2B pricing pages that day is in is your CDN blocking AI crawlers?.
How do you tell a real crawler from something using its name?
Check the request's IP against the vendor's published list or its reverse-DNS host: six of the seven pages link a list of the IP addresses their crawlers use; Meta's tells site owners to allow-list "the IP addresses (more secure)" and, as rendered, links no list. OpenAI publishes one JSON file per agent, four in all; Perplexity two; Anthropic one, adding that blocking its addresses "may not work correctly or persistently guarantee an opt-out, as doing so impedes our ability to read your robots.txt file". Google, Apple and Common Crawl also give a reverse-DNS host to check: googlebot.com or geo.googlebot.com, applebot.apple.com, and crawl.commoncrawl.org. Two pages say in as many words why this matters. Google: "The HTTP user agent string can be spoofed." Common Crawl: "we are aware of crawlers falsely identifying themselves as CCBot." The AI crawler user agents themselves are in the table; OpenAI's, Apple's and Google's pages each say the version number inside their strings changes over time.
How was this read?
With one plain fetch per page on 24 September 2026 between 19:07 and 19:08 UTC, a Chrome user-agent string and no JavaScript, from one machine in one location. Six pages answered 200 with their text in the HTML: OpenAI's Overview of OpenAI Crawlers, Perplexity Crawlers, Anthropic's help article on its crawlers, Google's list of common crawlers, Apple's About Applebot and Common Crawl's CCBot page. Meta's Web Crawlers page answered 400 to that fetch and was rendered in headless Chrome at 19:08 and again at 19:12 UTC, the second time with an English locale because the first render returned a page Meta marks as AI-translated into Spanish. No page showed a bot challenge, and none was bypassed. Four pages print a date and three print none; the table records each, and the two dated more than 90 days before our read, Anthropic's and Meta's, are marked with the gap. Two of the pages carry a banner pointing at an llms.txt index — Perplexity's, headed "For AI agents", tells the agent "Use this file to discover all available pages before exploring further" — and two more carry a button labelled Copy for LLM. We read those as data and followed none. Our own site serves an llms.txt at its root (it answered 200 on 24 September 2026), and what such a file has been shown to do is in does llms.txt actually work.
Each claim this post makes about our own system names its source in the same sentence, and you cannot check those from outside except the llms.txt, which you can request. Everything else is a quotation from a vendor's page, linked above and dated 24 September 2026, and a page can change the day after it is read. We did not observe any crawler, and nothing here says what one does.
Common questions
Does blocking Google-Extended affect Google AI Overviews?
Google's common-crawlers page, read 24 September 2026, does not mention AI Overviews or AI Mode. Google's page on AI features in Search, last updated 10 December 2025 and historical by our 90-day rule, says "robots.txt directives for Googlebot is the control" for those features, and points to Google-Extended "To limit AI training and grounding in some of Google's other systems". The common-crawlers page says is that Google-Extended controls whether crawled content "may be used for training future generations of Gemini models" and "for grounding" in Gemini Apps and Vertex AI, and that it "does not impact a site's inclusion in Google Search nor is it used as a ranking signal in Google Search".
Does Applebot-Extended crawl my site?
No, by Apple's page as read 24 September 2026: "Applebot-Extended does not crawl webpages." It is a robots.txt name that controls how data Applebot already crawled may be used for training Apple's foundation models, and "Webpages that disallow Applebot-Extended can still be included in search results".
Does blocking ClaudeBot stop Claude from reading my page when a user asks?
Not by itself, on Anthropic's page as read 24 September 2026: ClaudeBot is the agent that collects content "that could potentially contribute to their training", and Claude-User is the one that "may access websites" when a user asks a question. Disabling Claude-User "prevents our system from retrieving your content in response to a user query". The page says its bots honour robots.txt, and it prints no user-agent string for any of the three.
Which Meta user agent is the AI training crawler?
meta-externalagent, on Meta's page as rendered 24 September 2026: Meta-ExternalAgent "crawls the web for use cases such as training foundation AI models or improving products by indexing content directly". Meta-WebIndexer is the one Meta says "helps us cite and link to your content in Meta AI's responses" when allowed, and Meta-ExternalFetcher, which fetches "at a user's request", "may bypass robots.txt rules".
Does CCBot respect robots.txt?
Common Crawl's FAQ, linked from its CCBot page and read 24 September 2026, says CCBot is "an automated crawler, checking first the robots.txt" and that with a group naming CCBot and a Disallow line for the root path "our crawler will stop crawling your website". The CCBot page adds that Common Crawl is "aware of crawlers falsely identifying themselves as CCBot", and it links an IP list and a reverse-DNS host to check a request against. The user-agent string is CCBot/2.0 followed by a link to its FAQ.
These are our numbers. Yours are one audit away.
Seven pages, read on one day and quoted rather than summarised, every quote dated 24 September 2026 because a vendor can rewrite its crawler page without notice. Whether those crawlers reach your site is a question for your logs. Whether the engines they feed name you is a different one — 40 buyer prompts, you against up to three competitors that you name, as read on 19 September 2026 in our post on which engines we check — and that is what HeardOf does.
Is your CDN blocking AI crawlers? What Cloudflare's defaults do, per its own pages