The AI Crawler Directory
A plain-English reference for the bots that read the web on behalf of AI systems. For each crawler you'll find who runs it, what it's used for, its exact user-agent token, and how to allow or block it in robots.txt.
All 46 crawlers compared
Every token in one table, grouped by operator so each company's bots read as a set. "Respects robots.txt" reflects what the operator documents - not what independent measurement observes. Verification is the strongest published way to prove traffic is genuine.
| Crawler | Operator | Purpose | Respects robots.txt | Verification |
|---|---|---|---|---|
TongyiBotcommunityTongyiBot | Alibaba | Model training | Mostly | None published |
AI2BotAI2Bot | Allen Institute for AI | Model training | Mostly | None published |
AmazonbotAmazonbot | Amazon | Search & answers | Yes | rDNS + IP feed |
Amzn-SearchBotAmzn-SearchBot | Amazon | Search & answers | Yes | IP feed |
Amzn-UserAmzn-User | Amazon | User-triggered fetch | No (user-triggered) | IP feed |
bedrockbotbedrockbot | Amazon (AWS) | User-triggered fetch | Yes | None published |
Claude-CodecommunityClaude-Code | Anthropic | User-triggered fetch | Yes | None published |
Claude-SearchBotClaude-SearchBot | Anthropic | Search & answers | Yes | IP feed |
Claude-UserClaude-User | Anthropic | User-triggered fetch | Yes | IP feed |
ClaudeBotClaudeBot | Anthropic | Model training | Yes | IP feed |
ApplebotApplebot | Apple | Search & answers | Yes | rDNS + IP feed |
Applebot-ExtendedApplebot-Extended | Apple | Training opt-out token | Yes | None published |
BytespidercommunityBytespider | ByteDance | Model training | Inconsistent | None published |
TikTokSpidercommunityTikTokSpider | ByteDance | Model training | Inconsistent | None published |
Cohere Crawlercommunitycohere-training-data-crawler | Cohere | Model training | Mostly | None published |
cohere-aicommunitycohere-ai | Cohere | User-triggered fetch | Mostly | None published |
CCBotCCBot | Common Crawl Foundation | Model training | Yes | rDNS + IP feed |
DiffbotDiffbot | Diffbot | Model training | Mostly | None published |
DuckAssistBotDuckAssistBot | DuckDuckGo | Search & answers | Yes | IP feed |
ExaBotlegacyExaBot | Exa | Search & answers | Yes | Signed requests |
Gemini-Deep-ResearchcommunityGemini-Deep-Research | User-triggered fetch | No (user-triggered) | IP feed | |
Google-AgentGoogle-Agent | User-triggered fetch | No (user-triggered) | IP feed | |
Google-CloudVertexBotGoogle-CloudVertexBot | User-triggered fetch | Yes | IP feed | |
Google-ExtendedGoogle-Extended | Training opt-out token | Yes | None published | |
Google-NotebookLMlegacyGoogle-NotebookLM | User-triggered fetch | No (user-triggered) | IP feed | |
GoogleAgent-URLContextcommunityGoogleAgent-URLContext | User-triggered fetch | No (user-triggered) | IP feed | |
GooglebotGooglebot | Search & answers | Yes | rDNS + IP feed | |
GoogleOtherGoogleOther | General crawl | Yes | rDNS + IP feed | |
PanguBotcommunityPanguBot | Huawei | Model training | Mostly | None published |
PetalBotPetalBot | Huawei (Aspiegel) | Search & answers | Yes | rDNS + IP feed |
FacebookBotlegacyFacebookBot | Meta | Model training | Yes | None published |
Meta-ExternalAgentmeta-externalagent | Meta | Model training | Yes | None published |
Meta-ExternalFetchermeta-externalfetcher | Meta | User-triggered fetch | Mostly | None published |
Meta-WebIndexermeta-webindexer | Meta | Search & answers | Yes | None published |
BingbotBingbot | Microsoft | Search & answers | Yes | rDNS + IP feed |
MistralAI-UserMistralAI-User | Mistral AI | User-triggered fetch | Yes | IP feed |
Kimi-UsercommunityKimi-User | Moonshot AI | User-triggered fetch | Mostly | None published |
ChatGPT-UserChatGPT-User | OpenAI | User-triggered fetch | No (user-triggered) | IP feed |
GPTBotGPTBot | OpenAI | Model training | Yes | IP feed |
OAI-AdsBotOAI-AdsBot | OpenAI | Search & answers | Yes | IP feed |
OAI-SearchBotOAI-SearchBot | OpenAI | Search & answers | Yes | IP feed |
Perplexity-UserPerplexity-User | Perplexity | User-triggered fetch | No (user-triggered) | IP feed |
PerplexityBotPerplexityBot | Perplexity | Search & answers | Yes | IP feed |
Webzio-ExtendedWebzio-Extended | Webz.io | Training opt-out token | Yes | None published |
YouBotYouBot | You.com | Search & answers | Yes | None published |
ChatGLM-SpidercommunityChatGLM-Spider | Zhipu AI (attributed) | Model training | Mostly | None published |
Bots marked legacy use a deprecated or superseded token; community means the operator publishes little or no documentation, so the details come from crawler registries rather than vendor confirmation. Both are explained on the individual pages.
Browse each crawler in detail
One page per crawler: its full user-agent string, how to verify the traffic is genuine, which bots it gets confused with, and the exact robots.txt rules to allow or block it.
GPTBot
by OpenAI
OpenAI's training crawler. Allowing it lets your content be used to train future GPT models.
GPTBotOAI-SearchBot
by OpenAI
OpenAI's search crawler. Controls whether your pages can surface and be cited in ChatGPT's search features.
OAI-SearchBotChatGPT-User
by OpenAI
Fetches a page in real time when a ChatGPT user or agent follows a link. OpenAI states robots.txt may not apply to it.
ChatGPT-UserClaudeBot
by Anthropic
Anthropic's training crawler for Claude. One of the few AI crawlers that honors Crawl-delay.
ClaudeBotClaude-SearchBot
by Anthropic
Anthropic's search crawler. Controls whether your pages can be cited when Claude searches the web.
Claude-SearchBotClaude-User
by Anthropic
Fetches a page live when a Claude user asks Claude to read it. Notably, Anthropic says it respects do-not-crawl signals.
Claude-UserPerplexityBot
by Perplexity
Indexes pages so they can be cited in Perplexity's answers. Perplexity states it does not train foundation models.
PerplexityBotPerplexity-User
by Perplexity
Fetches a page in real time when answering a user's question needs it. Documented as generally ignoring robots.txt.
Perplexity-UserGoogle-Extended
by Google
A robots.txt token, not a crawler. Controls Gemini training use without touching Google Search.
Google-ExtendedGooglebot
by Google
Google's main search crawler - and the index behind AI Overviews. There is no separate AI Overviews bot.
GooglebotCCBot
by Common Crawl Foundation
Common Crawl's nonprofit crawler. Its open dataset is a starting point for a large share of LLM training pipelines.
CCBotBytespider
by ByteDance
ByteDance's training crawler. Widely documented as not honoring robots.txt, and notorious for crawl volume.
BytespiderAmazonbot
by Amazon
Amazon's general crawler - and the only one of its three that Amazon says may be used to train AI models.
AmazonbotApplebot-Extended
by Apple
A control token that opts you out of Apple's generative training. It does not crawl and cannot affect search inclusion.
Applebot-ExtendedMeta-ExternalAgent
by Meta
Meta's training crawler for Llama and Meta AI. Respects robots.txt, unlike two of its siblings.
meta-externalagentBingbot
by Microsoft
Microsoft's search crawler and the index behind Copilot. There is no separate Copilot crawler.
BingbotApplebot
by Apple
Apple's crawler for Siri and Spotlight. One of the few AI-relevant crawlers that fully renders JavaScript.
ApplebotMeta-ExternalFetcher
by Meta
Meta's user-triggered fetcher for Meta AI's agentic features. Meta states it may bypass robots.txt.
meta-externalfetcherMistralAI-User
by Mistral AI
Mistral's user-triggered fetcher. One of three separate Mistral tokens, cleanly split by purpose.
MistralAI-UserGoogle-CloudVertexBot
by Google
Crawls sites for Vertex AI Agents. Google documents it as crawling at the site owner's own request.
Google-CloudVertexBotAI2Bot
by Allen Institute for AI
A nonprofit research crawler for fully open models. Allowing it means your content lands in a public dataset.
AI2BotDuckAssistBot
by DuckDuckGo
Powers DuckAssist. DuckDuckGo states the data is never used to train AI models, and opting out takes 72 hours.
DuckAssistBotcohere-ai
by Cohere
Widely miscategorised as a training crawler - it actually retrieves pages for user-initiated prompts.
cohere-aiOAI-AdsBot
by OpenAI
OpenAI's ad-safety crawler. It reviews pages submitted as ChatGPT ads - not organic citation candidates.
OAI-AdsBotGoogleOther
by Google
Google's generic internal crawler. Google does not document a training purpose for it - despite common claims.
GoogleOtherGoogle-NotebookLM
by Google
Deprecated. Google now identifies Notebook source fetches as Google-GeminiNotebook - set rules for both.
Google-NotebookLMGoogle-Agent
by Google
Google's agentic fetcher - Gemini navigating and acting on a user's behalf. Documented to ignore robots.txt.
Google-AgentFacebookBot
by Meta
Meta's legacy training crawler, no longer listed on Meta's current crawler page. Superseded by meta-externalagent.
FacebookBotbedrockbot
by Amazon (AWS)
AWS's crawler for Bedrock knowledge bases. Its user-agent carries a per-crawler UUID, which trips exact-match rules.
bedrockbotPetalBot
by Huawei (Aspiegel)
Huawei's search crawler, operated by its Aspiegel subsidiary - which is why it verifies under aspiegel.com.
PetalBotPanguBot
by Huawei
Huawei's training crawler for its PanGu models. Far more thinly documented than Huawei's own PetalBot.
PanguBotCohere Crawler
by Cohere
Cohere's actual training crawler - the token to block if you want out of Cohere's training data.
cohere-training-data-crawlerDiffbot
by Diffbot
A commercial extraction crawler. Mass crawls honor robots.txt - but customer-requested single URLs may not.
DiffbotYouBot
by You.com
You.com's crawler. Unusually, one token covers both search indexing and LLM training - you cannot split them.
YouBotExaBot
by Exa
Exa's older token - the current one is ExaSearchBot, which is cryptographically signed rather than IP-verified.
ExaBotTikTokSpider
by ByteDance
ByteDance's second crawler token. Blocking Bytespider alone leaves this one running.
TikTokSpiderWebzio-Extended
by Webz.io
Webz.io's AI-training token. Unlike other -Extended tokens, this one belongs to a crawler that really does fetch.
Webzio-ExtendedKimi-User
by Moonshot AI
Moonshot AI's fetcher for the Kimi assistant - notable for long-context reading of entire documents.
Kimi-UserTongyiBot
by Alibaba
Alibaba's crawler for the Qwen model family - one of the most widely deployed open-weight model lines anywhere.
TongyiBotChatGLM-Spider
by Zhipu AI (attributed)
The least documented token in this directory - even its operator and purpose are unconfirmed attributions.
ChatGLM-SpiderMeta-WebIndexer
by Meta
Meta's search-index crawler. This is the token that decides whether Meta AI can cite you - not the training one.
meta-webindexerAmzn-SearchBot
by Amazon
Amazon's search crawler, documented as explicitly not crawling for generative AI training.
Amzn-SearchBotAmzn-User
by Amazon
Amazon's user-triggered fetcher. Amazon states it may not follow all robots.txt directives.
Amzn-UserGemini-Deep-Research
by Google
Fetches sources when a Gemini user runs Deep Research. Tracked by crawler registries but absent from Google's crawler docs.
Gemini-Deep-ResearchGoogleAgent-URLContext
by Google
Fetches URLs that Gemini API developers pass as context. Community-tracked; absent from Google's crawler docs.
GoogleAgent-URLContextClaude-Code
by Anthropic
The user-agent Claude Code sends when a developer asks it to fetch a URL. Its audience is developers, not consumers.
Claude-CodeReference
How to allow or block AI crawlers
AI crawlers are controlled the same way as any other bot: with robots.txt rules targeting each user-agent token. Add a block per bot you want to control, then a final catch-all if needed. Most reputable AI bots honor these rules - a few (noted on their pages) do not, and need server- or firewall-level blocking.
To allow a specific crawler:
User-agent: GPTBot Allow: /
To block a specific crawler:
User-agent: Bytespider Disallow: /
A common policy - opt out of AI training but stay visible in AI search and to traditional search engines:
User-agent: GPTBot Disallow: / User-agent: CCBot Disallow: / User-agent: Google-Extended Disallow: / User-agent: OAI-SearchBot Allow: / User-agent: PerplexityBot Allow: /
Note: Google-Extended and Applebot-Extended are training opt-out tokens - disallowing them does not remove you from Google Search, Siri, or Spotlight.
Not sure which bots can actually read your page?
Run any URL through the AI Crawl Checker to see exactly what these crawlers receive - content, status codes, structured data, and robots.txt rules - in seconds.
Check my site