Webzio-Extended
Webz.io's AI-training token. Unlike other -Extended tokens, this one belongs to a crawler that really does fetch.
Reference reviewed by Sona
| Operator | Webz.io |
|---|---|
| Powers | Webz.io datasets licensed to AI companies for training |
| Purpose | Training opt-out token |
| User-agent token | Webzio-Extended |
| Respects robots.txt | Yes |
Webz.io crawls the web - historically under the omgili and omgilibot user-agents - and licenses the resulting datasets to customers, AI companies buying training data among them. Webzio-Extended is the token that governs AI-training use of that data.
The -Extended suffix invites a wrong assumption. Google-Extended and Applebot-Extended send no traffic at all; they only annotate an existing crawl. Webz.io's arrangement is different: reporting on the transition from omgilibot describes webzio-extended as performing validation and then tagging collected data as usable or not usable for AI and ML training. So there is a crawler here, and it does fetch.
The practical consequence is that a complete opt-out needs more than one rule. Webz.io's guidance is to cover the old and new identifiers together - Omgilibot, Omgili, and webzio-extended - because rules naming only one of the three leave the others matching your wildcard group.
The leverage argument resembles Common Crawl's. Because Webz.io sells to many customers, a single block reduces your exposure across multiple downstream AI companies at once - with the difference that Webz.io's datasets are commercial and licensed rather than free and public, so there is an identifiable vendor and a contract behind each use.
How Webzio-Extended behaves
- Genuinely crawls, unlike the Google-Extended and Applebot-Extended tokens it resembles by name.
- Validates and tags collected data as usable or not usable for AI and ML training.
- Succeeds the omgili and omgilibot identifiers, which still appear in logs and still need rules.
- Feeds commercial licensed datasets, so exposure is to Webz.io's paying customers.
Allow Webzio-Extended
Your content can appear in commercial web datasets that many AI builders license and train on.
User-agent: Webzio-Extended Allow: /
Block Webzio-Extended
One decision that removes your content from training datasets resold to multiple AI companies - but write rules for the omgili tokens too.
User-agent: Webzio-Extended Disallow: /
How to verify Webzio-Extended
No published verification method
Webz.io publishes no IP range feed or reverse-DNS convention, so its crawler traffic cannot be authenticated. The complication specific to this vendor is identity spread across three names - Omgilibot, Omgili, and webzio-extended - so log analysis needs to recognize all of them as one operation before you can tell whether your rules are working.
Check an IP against this botCommonly confused with Webzio-Extended
Shares the -Extended naming but not the design: Google-Extended sends no traffic at all, while Webz.io's token belongs to a crawler that does fetch.
Both sell web data commercially. Webz.io licenses crawl datasets with training-usability tagging; Diffbot sells structured extractions and a knowledge graph.
Comparable one-rule leverage across many downstream models, but Webz.io's datasets are licensed and commercial where Common Crawl's are free and public.
Webzio-Extended FAQs
Is Webzio-Extended a control token like Google-Extended?
Not quite. Google-Extended and Applebot-Extended send no traffic and only govern usage. Webz.io's token is attached to a crawler that does fetch pages, validate them, and tag the data for AI-training usability.
Which tokens do I need to block Webz.io completely?
Three: Omgilibot, Omgili, and webzio-extended. Webz.io's own guidance is to cover the old and new identifiers together, since a rule naming only one leaves the others on your wildcard group.
Why does blocking Webz.io have outsized effect?
Because it sells datasets to many customers, so one block reduces your exposure across multiple downstream AI companies at once - similar leverage to blocking CCBot, but commercial and licensed rather than free and public.
What was omgilibot?
Webz.io's earlier crawler identity. The company introduced webzio-extended as part of a shift toward explicit AI-training tagging and opt-out, but the older identifiers still appear in logs.
Can Webzio-Extended read your page right now?
Test any URL and see exactly what AI crawlers receive.