← All guides

Guides

AI crawler user agents, verified: which GPTBot requests are real, and what each bot does

Published 7 Oct 2026

In short

A user agent proves nothing, because anyone can send "GPTBot". A request is really from OpenAI, Anthropic, Perplexity or Common Crawl when its IP address is inside the list that operator publishes; Google, Bing and Apple publish lists too and also document a reverse-DNS check, while Meta and ByteDance currently document neither for their AI crawlers. On our own site, all 16 GPTBot requests that failed the IP check were our own monitoring, which left 0 failures in 52.

Two ways to check a crawler, and which operators support each

The user-agent header is a claim the client makes about itself. Uptime monitors, SEO tools and scrapers all send "GPTBot" or "PerplexityBot" when it suits them, so counting user agents counts claims. Operators give you two ways to check the claim against something the client cannot fake.

  • IP list: the operator publishes the address ranges its crawler uses, as a JSON file. Check whether the request's source IP falls inside one of the ranges. OpenAI, Anthropic, Perplexity, Google, Microsoft (Bing), Apple and Common Crawl all publish one. Fetch the file again regularly: Microsoft asks you to refresh its list daily, because it can change at any time.
  • Reverse DNS, then forward DNS: look up the hostname for the request's IP, check it ends in the operator's domain, then resolve that hostname and confirm it returns the same IP. The second step matters, because anyone can set a reverse record for their own address. Google, Microsoft, Apple and Common Crawl document this method.
  • Neither: Meta's crawler page, read on 7 October 2026, tells you to allow-list its crawlers' IP addresses but publishes no list and no DNS method. We found no page from ByteDance documenting Bytespider at all.
  • Not every name in robots.txt is a crawler. Google-Extended and Applebot-Extended are robots.txt tokens only: they never appear as a user agent in your logs, and they control how content fetched by Googlebot or Applebot may be used.

For which tokens to allow or block, see how to audit your robots.txt and CDN for AI crawlers. This guide is the lookup table behind it.

OpenAI: GPTBot, OAI-SearchBot, ChatGPT-User

  • GPTBot. Purpose: training; OpenAI uses it to make its generative AI foundation models more useful and safe. robots.txt token: GPTBot; disallowing it tells OpenAI the site's content should not be used to train those models. IP list: openai.com/gptbot.json. Verify by: IP list.
  • OAI-SearchBot. Purpose: search index; it surfaces websites in ChatGPT's search features. robots.txt token: OAI-SearchBot; OpenAI says sites opted out of it are not shown in ChatGPT search answers, and that a robots.txt change takes about 24 hours to reach its systems. IP list: openai.com/searchbot.json. Verify by: IP list.
  • ChatGPT-User. Purpose: user-triggered fetch, for certain user actions in ChatGPT and Custom GPTs; OpenAI says it does not crawl the web automatically. robots.txt: OpenAI says that because a user initiated the action, robots.txt rules may not apply. IP list: openai.com/chatgpt-user.json. Verify by: IP list.
  • OAI-AdsBot. Purpose: checks the safety of web pages submitted as ads on ChatGPT. IP list: openai.com/adsbot.json. The page does not say how it treats robots.txt.
  • OpenAI states that each robots.txt setting is independent of the others: you can allow OAI-SearchBot and disallow GPTBot.

Anthropic: ClaudeBot, Claude-SearchBot, Claude-User

  • ClaudeBot. Purpose: training; it collects web content that could contribute to training Anthropic's models. robots.txt token: ClaudeBot; disallowing it signals that the site's future material should be excluded from training data.
  • Claude-SearchBot. Purpose: search index; Anthropic says it navigates the web to improve the quality of search results for Claude's users. robots.txt token: Claude-SearchBot.
  • Claude-User. Purpose: user-triggered fetch; when someone asks Claude a question, it may access websites with this agent. robots.txt token: Claude-User; Anthropic says disabling it prevents its system from retrieving your content in response to a user's query.
  • robots.txt, all three: Anthropic says its bots honour robots.txt directives, and it supports the non-standard Crawl-delay extension. Unlike OpenAI and Perplexity, it makes no exception for the user-triggered agent.
  • IP list, all three: one shared file, claude.com/crawling/bots.json. Anthropic says a source IP on that list means the crawler is coming from Anthropic. The file does not say which of the three bots uses which range. Verify by: IP list. Anthropic also warns that blocking its IP addresses is not a reliable opt-out; use robots.txt.

Perplexity: PerplexityBot, Perplexity-User

  • PerplexityBot. Purpose: search index; it surfaces and links websites in Perplexity's search results. Perplexity says it is not used to crawl content for AI foundation models. robots.txt token: PerplexityBot. IP list: perplexity.com/perplexitybot.json. Verify by: IP list.
  • Perplexity-User. Purpose: user-triggered fetch, supporting user actions within Perplexity; not used for web crawling or for training. robots.txt: Perplexity says that since a user requested the fetch, it generally ignores robots.txt rules. IP list: perplexity.com/perplexity-user.json. Verify by: IP list.

Google: Googlebot, Google-Extended, and AI Overviews

  • Googlebot. Purpose: search index, for Google Search including all of its search features. robots.txt token: Googlebot. Google says its common crawlers always obey robots.txt when crawling automatically. Verify by: reverse DNS to googlebot.com, google.com or googleusercontent.com followed by a matching forward lookup, or the IP list developers.google.com/static/crawling/ipranges/common-crawlers.json.
  • AI Overviews and AI Mode have no crawler of their own. Google's page on AI features says a page must be indexed and eligible to be shown in Google Search with a snippet, and that there are no additional technical requirements. The controls Google points to for limiting what those features show are ordinary Search controls: nosnippet, data-nosnippet, max-snippet and noindex.
  • Google-Extended. Not a crawler: Google says it has no separate HTTP user agent, so it never appears in your logs. It is a robots.txt token that controls whether content Google crawls may be used to train future Gemini models (Gemini Apps, Vertex AI API for Gemini) and for grounding in Gemini Apps and Grounding with Google Search on Vertex AI. Google says it does not affect inclusion in Google Search and is not a ranking signal.
  • Google publishes separate IP files for special crawlers and for user-triggered fetchers, in the same folder as the common-crawlers file. A request from a user-triggered fetcher is not Googlebot, even when it comes from Google.

Microsoft: Bingbot and Copilot

  • Bingbot. Purpose: search index, Bing's standard crawler. Microsoft lists no separate crawler for Copilot: its webmaster guidelines say Bing and Copilot search experiences rely on the same crawling, indexing and ranking foundation, and list blocking Bingbot in robots.txt among the things to avoid. robots.txt token: bingbot.
  • Controlling Copilot is done with meta directives on the page rather than a robots.txt token. Microsoft says NOARCHIVE keeps content out of Copilot responses and grounding results, and NOCACHE limits Copilot to the URL, title and snippet.
  • Verify by: reverse DNS to a name ending in search.msn.com followed by a forward lookup that returns the same IP; Microsoft's Verify Bingbot tool; or the IP list bing.com/toolbox/bingbot.json, which Microsoft asks you to refresh daily.

Apple: Applebot and Applebot-Extended

  • Applebot. Purpose: search index; Apple says its data powers features such as the search technology in Spotlight, Siri and Safari, and may also be used to help train Apple foundation models powering generative AI features. robots.txt token: Applebot; Apple says it respects robots.txt directives aimed at Applebot, follows Googlebot instructions when the file mentions Googlebot but not Applebot, and does not follow crawl-delay. Verify by: reverse DNS in applebot.apple.com with a matching forward lookup, or the IP list search.developer.apple.com/applebot.json.
  • Applebot-Extended. Not a crawler: Apple says it does not crawl webpages. It is a robots.txt token that decides whether content Applebot already fetched may train Apple's general-purpose foundation models. Pages that disallow it can still be included in search results.

Meta, ByteDance and Common Crawl

  • Meta-ExternalAgent. Operator: Meta. Purpose: training and indexing; Meta says it crawls for use cases such as training foundation AI models or improving products by indexing content directly. robots.txt token: meta-externalagent. IP list: none on Meta's crawler page as read on 7 October 2026.
  • Meta-ExternalFetcher. Operator: Meta. Purpose: user-triggered fetch of individual links, including helping AI complete tasks for users. robots.txt: Meta says it may bypass robots.txt because the fetch was requested by a user. IP list: none published on the same page.
  • Bytespider. Operator: ByteDance. We found no page from ByteDance documenting its purpose, its robots.txt behaviour or its addresses, so there is no vendor statement to put here. Everything circulating about it comes from third parties.
  • CCBot. Operator: Common Crawl, a non-profit that produces and maintains an open repository of web crawl data. robots.txt token: CCBot; Common Crawl documents Disallow for it. IP list: index.commoncrawl.org/ccbot.json. Verify by: IP list, or reverse DNS (IPv4 addresses resolve under crawl.commoncrawl.org; Common Crawl says IPv6 has no reverse DNS yet).

What our own log showed

One site, ours: promvia.app, 26 August to 5 October 2026. We checked every request whose user agent claimed GPTBot, OAI-SearchBot, ChatGPT-User or PerplexityBot against the operator's published IP list. Read as raw failure rates, the result looks like an impostor problem.

  • GPTBot: 16 of 68 requests failed verification (24%). All 16 were ours. The rest: 0 of 52.
  • PerplexityBot: 9 of 34 failed (26%). All 9 were ours. The rest: 0 of 25.
  • OAI-SearchBot: 7 of 390 failed (1.8%). All 7 were ours. The rest: 0 of 383.
  • ChatGPT-User: 17 of 610 failed (2.8%). 7 were ours; 10 of 603 (1.7%) we cannot explain.
  • Perplexity-User never appeared in the log.

"Ours" is Promvia's own weekly crawler-access check, which requests our homepage once with each crawler's user agent, and one address we used for manual tests. If you run a monitor or an SEO tool that sends bot user agents, subtract it before you count fakes.

Verification also changed what we could say about the index crawlers. Counting only verified requests, OAI-SearchBot had fetched the page first in 116 of 129 cases where ChatGPT's API answer cited a promvia.app URL, and PerplexityBot in 174 of 174 cases for Perplexity's. The questions all name our brand, so this is one site's log, not a rule. The full join, with the rows and the script, is in joining server logs to AI citations.

How Promvia checks crawlers, and where it cannot

  • Where the ranges come from: once a week, on Monday, Promvia fetches the operators' IP lists through an open-source aggregate that copies each operator's own published file, and caches them per crawler.
  • What "Verified" means: the request reached us from your CDN log connector or your token-authenticated edge forwarder, AND its source IP is inside a cached range for the crawler its user agent names. Either condition alone is a claim, so the hit stays "Detected".
  • What stays "Detected": Bingbot, Bytespider and Meta's agents. We do not load Bing's list yet, and ByteDance and Meta publish no crawler list. Anthropic's single shared list verifies all three of its agents, ClaudeBot, Claude-SearchBot and Claude-User, because Anthropic publishes it for all three. Googlebot is not counted as an AI crawler at all.
  • Meta is not verified at all. The only list available is Meta's network-wide address list, which covers every address Meta routes, so a match would only prove a request came from Meta's network, not from its crawler. We would rather show "Detected" than a "Verified" that means something else.
  • Ranges follow the operators' lists both ways. When a crawler's list comes back in full, the refresh replaces our copy, so a range the operator withdrew stops verifying. When a list fails or comes back empty, we keep the previous copy rather than wipe it, so real crawlers read as "Detected" at worst, never as a false "Verified".

Two related guides cover what this table does not: what to put in robots.txt, in the robots.txt and CDN audit, and whether llms.txt is read at all, in the honest guide to llms.txt. For a site that refuses every crawler in robots.txt and grants access through licensing deals instead, see Reddit and ChatGPT's citations.

Frequently asked questions

How do I know a GPTBot request is really from OpenAI?

Check its source IP against openai.com/gptbot.json, the list OpenAI publishes for GPTBot. OpenAI documents no reverse-DNS method, so the IP list is the check. The user agent alone proves nothing.

Is Google-Extended a crawler I can find in my logs?

No. Google says Google-Extended has no HTTP user agent of its own. It is a robots.txt token that controls whether content Googlebot already crawls may train Gemini models and ground Gemini answers. It does not affect Google Search inclusion or ranking.

Do AI crawlers respect robots.txt?

For the training and index crawlers, OpenAI, Anthropic, Perplexity, Google, Microsoft, Apple and Common Crawl all document robots.txt controls. The user-triggered fetchers differ. OpenAI says robots.txt may not apply to ChatGPT-User, Perplexity says Perplexity-User generally ignores it, and Meta says Meta-ExternalFetcher may bypass it. Anthropic says all three of its bots, including Claude-User, honour it.

Which AI crawlers publish no IP list?

As of 7 October 2026, Meta publishes none for Meta-ExternalAgent or Meta-ExternalFetcher on its crawler page, and we found no ByteDance documentation for Bytespider. Anthropic publishes one shared list for all three Claude agents rather than one per agent.

See where you stand today

Run your questions across up to 7 AI visibility surfaces and get your baseline — small runs often finish in about a minute. Free plan, no credit card.

Start free →Try the free AI Traffic Checker

Keep reading

How to audit your robots.txt (and CDN) for AI crawlers →Joining server logs to AI citations: what ChatGPT's and Perplexity's crawlers did before they cited us →llms.txt: what it is, and what it actually does (an honest guide) →