Cohere Crawler
Cohere's actual training crawler - the token to block if you want out of Cohere's training data.
Reference reviewed by Sona
The operator publishes little or no documentation for this token. Details here come from community crawler registries and observed behavior, not vendor confirmation.
| Operator | Cohere |
|---|---|
| Powers | Cohere enterprise LLM training |
| Purpose | Model training |
| User-agent token | cohere-training-data-crawler |
| Respects robots.txt | Mostly |
cohere-training-data-crawler downloads training data for the large language models behind Cohere's enterprise AI products. Despite the awkward, self-describing name, this - not cohere-ai - is the token that governs whether your content enters Cohere's training corpus.
The pairing is the thing to get right. cohere-ai is documented in community registries as serving user-initiated prompts, so a training opt-out that names only cohere-ai has blocked live retrieval and left the training crawl running. Set rules for both tokens if your intent is a complete Cohere opt-out.
Cohere publishes no crawler documentation page for either token - no purpose statement, no compliance policy, no IP feed. Community registries and behavioral reports describe standard robots.txt compliance, which is plausible for an enterprise vendor with reputational exposure to enterprise buyers, but it is not a vendor commitment and should not be described as one.
Cohere's positioning shapes what allowing this actually means. Its models are built for business search, retrieval, and generation deployments rather than a consumer assistant, so your content would be influencing enterprise tooling rather than public-facing chat answers.
How Cohere Crawler behaves
- Community-documented as downloading training data for Cohere's enterprise LLMs.
- Compliance is behaviorally reported rather than vendor-confirmed - Cohere publishes no policy.
- No IP feed, reverse-DNS convention, or crawler documentation page.
- Needs pairing with a cohere-ai rule for a complete Cohere opt-out.
Allow Cohere Crawler
Your content can be represented in Cohere-powered enterprise search and generation products.
User-agent: cohere-training-data-crawler Allow: /
Block Cohere Crawler
This is the token that actually keeps your content out of Cohere's training data - and it needs to be paired with a cohere-ai rule for full coverage.
User-agent: cohere-training-data-crawler Disallow: /
How to verify Cohere Crawler
No published verification method
Cohere publishes no IP range feed, no reverse-DNS convention, and no crawler documentation, so this traffic cannot be authenticated. The token is distinctive enough that impersonation is less attractive than with shorter, better-known names, but that is a weak comfort rather than a verification method. Edge rules must match the user-agent on trust.
Check an IP against this botCommonly confused with Cohere Crawler
Cohere Crawler FAQs
Which Cohere token do I block to opt out of training?
cohere-training-data-crawler. The shorter cohere-ai token is documented in community registries as serving user-initiated prompts, so blocking only that one leaves the training crawl unaffected.
Does this crawler respect robots.txt?
Reportedly yes, but Cohere has not confirmed it. There is no crawler documentation page, no published compliance policy, and no IP feed - everything known comes from community registries and behavioral observation.
Do I need rules for both Cohere tokens?
Yes, if you want complete coverage. They do different jobs - one collects training data, the other serves live prompts - and each reads its own robots.txt group.
What does Cohere use the training data for?
Training the language models behind its enterprise products, which target business search, retrieval, and generation rather than a consumer chat assistant.
Can Cohere Crawler read your page right now?
Test any URL and see exactly what AI crawlers receive.