Sona

Diffbot

A commercial extraction crawler - turns your pages into structured data and a knowledge graph that AI companies license.

Reference reviewed by Sona

OperatorDiffbot
PowersStructured web data and knowledge graph resold to AI builders
PurposeModel training
User-agent tokenDiffbot
Respects robots.txtMostly

Diffbot crawls and machine-reads web pages into structured records and a large knowledge graph. Its customers - including AI companies - license this data for training, grounding, and enrichment, so allowing Diffbot indirectly feeds many downstream products.

Diffbot respects robots.txt by default, but as a commercial service its crawls can be configured per customer, so compliance is best described as default-on rather than absolute.

Full user-agent string

Mozilla/5.0 (compatible; Diffbot/2.0; +http://www.diffbot.com)

Allow Diffbot

Accurate structured data about your site propagates into the many products and AI systems built on Diffbot's knowledge graph.

User-agent: Diffbot
Allow: /

Block Diffbot

You don't want your content extracted and resold as structured data to third parties.

User-agent: Diffbot
Disallow: /

Can Diffbot read your page right now?

Test any URL and see exactly what AI crawlers receive.

Check my site