Sona

AI2Bot

The Allen Institute's crawler, gathering training data for fully open models like OLMo.

Reference reviewed by Sona

OperatorAllen Institute for AI
PowersOLMo and other open-model training
PurposeModel training
User-agent tokenAI2Bot
Respects robots.txtYes

AI2Bot collects publicly available web content for the Allen Institute for AI, whose OLMo family are fully open models - open weights, open data, and open training code.

AI2Bot publishes its user-agent and respects robots.txt. Because OLMo's training data is released openly, allowing AI2Bot also means your content appears in a public research corpus.

Full user-agent string

Mozilla/5.0 (compatible) AI2Bot (+https://www.allenai.org/crawler)

Allow AI2Bot

Your content supports open AI research and is represented in openly released models and datasets.

User-agent: AI2Bot
Allow: /

Block AI2Bot

You don't want your content in publicly released training corpora.

User-agent: AI2Bot
Disallow: /

Can AI2Bot read your page right now?

Test any URL and see exactly what AI crawlers receive.

Check my site