New: monitor which AI crawlers actually visit your site with Sona Agent Analytics  |  See the platform →

The AI Crawler Directory

A plain-English reference for the bots that read the web on behalf of AI systems. For each crawler you'll find who runs it, what it's used for, its exact user-agent token, and how to allow or block it in robots.txt.

All 46 crawlers compared

Every token in one table, grouped by operator so each company's bots read as a set. "Respects robots.txt" reflects what the operator documents - not what independent measurement observes. Verification is the strongest published way to prove traffic is genuine.

CrawlerOperatorPurposeRespects robots.txtVerification
TongyiBotcommunityTongyiBotAlibabaModel trainingMostlyNone published
AI2BotAI2BotAllen Institute for AIModel trainingMostlyNone published
AmazonbotAmazonbotAmazonSearch & answersYesrDNS + IP feed
Amzn-SearchBotAmzn-SearchBotAmazonSearch & answersYesIP feed
Amzn-UserAmzn-UserAmazonUser-triggered fetchNo (user-triggered)IP feed
bedrockbotbedrockbotAmazon (AWS)User-triggered fetchYesNone published
Claude-CodecommunityClaude-CodeAnthropicUser-triggered fetchYesNone published
Claude-SearchBotClaude-SearchBotAnthropicSearch & answersYesIP feed
Claude-UserClaude-UserAnthropicUser-triggered fetchYesIP feed
ClaudeBotClaudeBotAnthropicModel trainingYesIP feed
ApplebotApplebotAppleSearch & answersYesrDNS + IP feed
Applebot-ExtendedApplebot-ExtendedAppleTraining opt-out tokenYesNone published
BytespidercommunityBytespiderByteDanceModel trainingInconsistentNone published
TikTokSpidercommunityTikTokSpiderByteDanceModel trainingInconsistentNone published
Cohere Crawlercommunitycohere-training-data-crawlerCohereModel trainingMostlyNone published
cohere-aicommunitycohere-aiCohereUser-triggered fetchMostlyNone published
CCBotCCBotCommon Crawl FoundationModel trainingYesrDNS + IP feed
DiffbotDiffbotDiffbotModel trainingMostlyNone published
DuckAssistBotDuckAssistBotDuckDuckGoSearch & answersYesIP feed
ExaBotlegacyExaBotExaSearch & answersYesSigned requests
Gemini-Deep-ResearchcommunityGemini-Deep-ResearchGoogleUser-triggered fetchNo (user-triggered)IP feed
Google-AgentGoogle-AgentGoogleUser-triggered fetchNo (user-triggered)IP feed
Google-CloudVertexBotGoogle-CloudVertexBotGoogleUser-triggered fetchYesIP feed
Google-ExtendedGoogle-ExtendedGoogleTraining opt-out tokenYesNone published
Google-NotebookLMlegacyGoogle-NotebookLMGoogleUser-triggered fetchNo (user-triggered)IP feed
GoogleAgent-URLContextcommunityGoogleAgent-URLContextGoogleUser-triggered fetchNo (user-triggered)IP feed
GooglebotGooglebotGoogleSearch & answersYesrDNS + IP feed
GoogleOtherGoogleOtherGoogleGeneral crawlYesrDNS + IP feed
PanguBotcommunityPanguBotHuaweiModel trainingMostlyNone published
PetalBotPetalBotHuawei (Aspiegel)Search & answersYesrDNS + IP feed
FacebookBotlegacyFacebookBotMetaModel trainingYesNone published
Meta-ExternalAgentmeta-externalagentMetaModel trainingYesNone published
Meta-ExternalFetchermeta-externalfetcherMetaUser-triggered fetchMostlyNone published
Meta-WebIndexermeta-webindexerMetaSearch & answersYesNone published
BingbotBingbotMicrosoftSearch & answersYesrDNS + IP feed
MistralAI-UserMistralAI-UserMistral AIUser-triggered fetchYesIP feed
Kimi-UsercommunityKimi-UserMoonshot AIUser-triggered fetchMostlyNone published
ChatGPT-UserChatGPT-UserOpenAIUser-triggered fetchNo (user-triggered)IP feed
GPTBotGPTBotOpenAIModel trainingYesIP feed
OAI-AdsBotOAI-AdsBotOpenAISearch & answersYesIP feed
OAI-SearchBotOAI-SearchBotOpenAISearch & answersYesIP feed
Perplexity-UserPerplexity-UserPerplexityUser-triggered fetchNo (user-triggered)IP feed
PerplexityBotPerplexityBotPerplexitySearch & answersYesIP feed
Webzio-ExtendedWebzio-ExtendedWebz.ioTraining opt-out tokenYesNone published
YouBotYouBotYou.comSearch & answersYesNone published
ChatGLM-SpidercommunityChatGLM-SpiderZhipu AI (attributed)Model trainingMostlyNone published

Bots marked legacy use a deprecated or superseded token; community means the operator publishes little or no documentation, so the details come from crawler registries rather than vendor confirmation. Both are explained on the individual pages.

Browse each crawler in detail

One page per crawler: its full user-agent string, how to verify the traffic is genuine, which bots it gets confused with, and the exact robots.txt rules to allow or block it.

Model training

GPTBot

by OpenAI

OpenAI's training crawler. Allowing it lets your content be used to train future GPT models.

GPTBot
Search & answers

OAI-SearchBot

by OpenAI

OpenAI's search crawler. Controls whether your pages can surface and be cited in ChatGPT's search features.

OAI-SearchBot
User-triggered fetch

ChatGPT-User

by OpenAI

Fetches a page in real time when a ChatGPT user or agent follows a link. OpenAI states robots.txt may not apply to it.

ChatGPT-User
Model training

ClaudeBot

by Anthropic

Anthropic's training crawler for Claude. One of the few AI crawlers that honors Crawl-delay.

ClaudeBot
Search & answers

Claude-SearchBot

by Anthropic

Anthropic's search crawler. Controls whether your pages can be cited when Claude searches the web.

Claude-SearchBot
User-triggered fetch

Claude-User

by Anthropic

Fetches a page live when a Claude user asks Claude to read it. Notably, Anthropic says it respects do-not-crawl signals.

Claude-User
Search & answers

PerplexityBot

by Perplexity

Indexes pages so they can be cited in Perplexity's answers. Perplexity states it does not train foundation models.

PerplexityBot
User-triggered fetch

Perplexity-User

by Perplexity

Fetches a page in real time when answering a user's question needs it. Documented as generally ignoring robots.txt.

Perplexity-User
Training opt-out token

Google-Extended

by Google

A robots.txt token, not a crawler. Controls Gemini training use without touching Google Search.

Google-Extended
Search & answers

Googlebot

by Google

Google's main search crawler - and the index behind AI Overviews. There is no separate AI Overviews bot.

Googlebot
Model training

CCBot

by Common Crawl Foundation

Common Crawl's nonprofit crawler. Its open dataset is a starting point for a large share of LLM training pipelines.

CCBot
Model training

Bytespider

by ByteDance

ByteDance's training crawler. Widely documented as not honoring robots.txt, and notorious for crawl volume.

Bytespider
Search & answers

Amazonbot

by Amazon

Amazon's general crawler - and the only one of its three that Amazon says may be used to train AI models.

Amazonbot
Training opt-out token

Applebot-Extended

by Apple

A control token that opts you out of Apple's generative training. It does not crawl and cannot affect search inclusion.

Applebot-Extended
Model training

Meta-ExternalAgent

by Meta

Meta's training crawler for Llama and Meta AI. Respects robots.txt, unlike two of its siblings.

meta-externalagent
Search & answers

Bingbot

by Microsoft

Microsoft's search crawler and the index behind Copilot. There is no separate Copilot crawler.

Bingbot
Search & answers

Applebot

by Apple

Apple's crawler for Siri and Spotlight. One of the few AI-relevant crawlers that fully renders JavaScript.

Applebot
User-triggered fetch

Meta-ExternalFetcher

by Meta

Meta's user-triggered fetcher for Meta AI's agentic features. Meta states it may bypass robots.txt.

meta-externalfetcher
User-triggered fetch

MistralAI-User

by Mistral AI

Mistral's user-triggered fetcher. One of three separate Mistral tokens, cleanly split by purpose.

MistralAI-User
User-triggered fetch

Google-CloudVertexBot

by Google

Crawls sites for Vertex AI Agents. Google documents it as crawling at the site owner's own request.

Google-CloudVertexBot
Model training

AI2Bot

by Allen Institute for AI

A nonprofit research crawler for fully open models. Allowing it means your content lands in a public dataset.

AI2Bot
Search & answers

DuckAssistBot

by DuckDuckGo

Powers DuckAssist. DuckDuckGo states the data is never used to train AI models, and opting out takes 72 hours.

DuckAssistBot
User-triggered fetch

cohere-ai

by Cohere

Widely miscategorised as a training crawler - it actually retrieves pages for user-initiated prompts.

cohere-ai
Search & answers

OAI-AdsBot

by OpenAI

OpenAI's ad-safety crawler. It reviews pages submitted as ChatGPT ads - not organic citation candidates.

OAI-AdsBot
General crawl

GoogleOther

by Google

Google's generic internal crawler. Google does not document a training purpose for it - despite common claims.

GoogleOther
User-triggered fetch

Google-NotebookLM

by Google

Deprecated. Google now identifies Notebook source fetches as Google-GeminiNotebook - set rules for both.

Google-NotebookLM
User-triggered fetch

Google-Agent

by Google

Google's agentic fetcher - Gemini navigating and acting on a user's behalf. Documented to ignore robots.txt.

Google-Agent
Model training

FacebookBot

by Meta

Meta's legacy training crawler, no longer listed on Meta's current crawler page. Superseded by meta-externalagent.

FacebookBot
User-triggered fetch

bedrockbot

by Amazon (AWS)

AWS's crawler for Bedrock knowledge bases. Its user-agent carries a per-crawler UUID, which trips exact-match rules.

bedrockbot
Search & answers

PetalBot

by Huawei (Aspiegel)

Huawei's search crawler, operated by its Aspiegel subsidiary - which is why it verifies under aspiegel.com.

PetalBot
Model training

PanguBot

by Huawei

Huawei's training crawler for its PanGu models. Far more thinly documented than Huawei's own PetalBot.

PanguBot
Model training

Cohere Crawler

by Cohere

Cohere's actual training crawler - the token to block if you want out of Cohere's training data.

cohere-training-data-crawler
Model training

Diffbot

by Diffbot

A commercial extraction crawler. Mass crawls honor robots.txt - but customer-requested single URLs may not.

Diffbot
Search & answers

YouBot

by You.com

You.com's crawler. Unusually, one token covers both search indexing and LLM training - you cannot split them.

YouBot
Search & answers

ExaBot

by Exa

Exa's older token - the current one is ExaSearchBot, which is cryptographically signed rather than IP-verified.

ExaBot
Model training

TikTokSpider

by ByteDance

ByteDance's second crawler token. Blocking Bytespider alone leaves this one running.

TikTokSpider
Training opt-out token

Webzio-Extended

by Webz.io

Webz.io's AI-training token. Unlike other -Extended tokens, this one belongs to a crawler that really does fetch.

Webzio-Extended
User-triggered fetch

Kimi-User

by Moonshot AI

Moonshot AI's fetcher for the Kimi assistant - notable for long-context reading of entire documents.

Kimi-User
Model training

TongyiBot

by Alibaba

Alibaba's crawler for the Qwen model family - one of the most widely deployed open-weight model lines anywhere.

TongyiBot
Model training

ChatGLM-Spider

by Zhipu AI (attributed)

The least documented token in this directory - even its operator and purpose are unconfirmed attributions.

ChatGLM-Spider
Search & answers

Meta-WebIndexer

by Meta

Meta's search-index crawler. This is the token that decides whether Meta AI can cite you - not the training one.

meta-webindexer
Search & answers

Amzn-SearchBot

by Amazon

Amazon's search crawler, documented as explicitly not crawling for generative AI training.

Amzn-SearchBot
User-triggered fetch

Amzn-User

by Amazon

Amazon's user-triggered fetcher. Amazon states it may not follow all robots.txt directives.

Amzn-User
User-triggered fetch

Gemini-Deep-Research

by Google

Fetches sources when a Gemini user runs Deep Research. Tracked by crawler registries but absent from Google's crawler docs.

Gemini-Deep-Research
User-triggered fetch

GoogleAgent-URLContext

by Google

Fetches URLs that Gemini API developers pass as context. Community-tracked; absent from Google's crawler docs.

GoogleAgent-URLContext
User-triggered fetch

Claude-Code

by Anthropic

The user-agent Claude Code sends when a developer asks it to fetch a URL. Its audience is developers, not consumers.

Claude-Code

Reference

How to allow or block AI crawlers

AI crawlers are controlled the same way as any other bot: with robots.txt rules targeting each user-agent token. Add a block per bot you want to control, then a final catch-all if needed. Most reputable AI bots honor these rules - a few (noted on their pages) do not, and need server- or firewall-level blocking.

To allow a specific crawler:

User-agent: GPTBot
Allow: /

To block a specific crawler:

User-agent: Bytespider
Disallow: /

A common policy - opt out of AI training but stay visible in AI search and to traditional search engines:

User-agent: GPTBot
Disallow: /

User-agent: CCBot
Disallow: /

User-agent: Google-Extended
Disallow: /

User-agent: OAI-SearchBot
Allow: /

User-agent: PerplexityBot
Allow: /

Note: Google-Extended and Applebot-Extended are training opt-out tokens - disallowing them does not remove you from Google Search, Siri, or Spotlight.

Not sure which bots can actually read your page?

Run any URL through the AI Crawl Checker to see exactly what these crawlers receive - content, status codes, structured data, and robots.txt rules - in seconds.

Check my site