Cohere Crawler
Cohere's dedicated training-data crawler - the current token for keeping content in or out of Cohere's enterprise models.
Reference reviewed by Sona
| Operator | Cohere |
|---|---|
| Powers | Cohere enterprise LLM training |
| Purpose | Model training |
| User-agent token | cohere-training-data-crawler |
| Respects robots.txt | Mostly |
cohere-training-data-crawler is the token associated with Cohere's collection of web content for training its enterprise language models - the successor to the older, vaguer cohere-ai token.
Cohere publishes no dedicated crawler documentation page; behavioral reports describe standard robots.txt compliance, but the vendor has not formally confirmed it. Set rules for both this token and cohere-ai for full coverage.
Allow Cohere Crawler
Your content can be represented in Cohere-powered enterprise AI products used in business search and generation.
User-agent: cohere-training-data-crawler Allow: /
Block Cohere Crawler
Keep your content out of Cohere's training data.
User-agent: cohere-training-data-crawler Disallow: /
Can Cohere Crawler read your page right now?
Test any URL and see exactly what AI crawlers receive.