Diffbot
A commercial extraction crawler - turns your pages into structured data and a knowledge graph that AI companies license.
Reference reviewed by Sona
| Operator | Diffbot |
|---|---|
| Powers | Structured web data and knowledge graph resold to AI builders |
| Purpose | Model training |
| User-agent token | Diffbot |
| Respects robots.txt | Mostly |
Diffbot crawls and machine-reads web pages into structured records and a large knowledge graph. Its customers - including AI companies - license this data for training, grounding, and enrichment, so allowing Diffbot indirectly feeds many downstream products.
Diffbot respects robots.txt by default, but as a commercial service its crawls can be configured per customer, so compliance is best described as default-on rather than absolute.
Full user-agent string
Mozilla/5.0 (compatible; Diffbot/2.0; +http://www.diffbot.com)
Allow Diffbot
Accurate structured data about your site propagates into the many products and AI systems built on Diffbot's knowledge graph.
User-agent: Diffbot Allow: /
Block Diffbot
You don't want your content extracted and resold as structured data to third parties.
User-agent: Diffbot Disallow: /
Can Diffbot read your page right now?
Test any URL and see exactly what AI crawlers receive.