AI2Bot
The Allen Institute's crawler, gathering training data for fully open models like OLMo.
Reference reviewed by Sona
| Operator | Allen Institute for AI |
|---|---|
| Powers | OLMo and other open-model training |
| Purpose | Model training |
| User-agent token | AI2Bot |
| Respects robots.txt | Yes |
AI2Bot collects publicly available web content for the Allen Institute for AI, whose OLMo family are fully open models - open weights, open data, and open training code.
AI2Bot publishes its user-agent and respects robots.txt. Because OLMo's training data is released openly, allowing AI2Bot also means your content appears in a public research corpus.
Full user-agent string
Mozilla/5.0 (compatible) AI2Bot (+https://www.allenai.org/crawler)
Allow AI2Bot
Your content supports open AI research and is represented in openly released models and datasets.
User-agent: AI2Bot Allow: /
Block AI2Bot
You don't want your content in publicly released training corpora.
User-agent: AI2Bot Disallow: /
Can AI2Bot read your page right now?
Test any URL and see exactly what AI crawlers receive.