PanguBot
Huawei's training crawler for its PanGu models. Far more thinly documented than Huawei's own PetalBot.
Reference reviewed by Sona
The operator publishes little or no documentation for this token. Details here come from community crawler registries and observed behavior, not vendor confirmation.
| Operator | Huawei |
|---|---|
| Powers | PanGu large language model training |
| Purpose | Model training |
| User-agent token | PanguBot |
| Respects robots.txt | Mostly |
PanguBot gathers web content used to train Huawei's PanGu family of multimodal large language models, which power AI features across Huawei's cloud services and device products. PanGu is a serious model line - it is Huawei's answer to GPT and Gemini, deployed heavily in Chinese enterprise and government contexts.
The documentation gap next to PetalBot is striking, given both are Huawei. PetalBot has a webmaster portal, an operator page, and a documented reverse-DNS verification domain. PanguBot has none of that: no crawler documentation page, no IP feed, no verification mechanism. Everything here comes from community crawler registries, which list its robots.txt compliance as unclear.
In practice, treat compliance as expected but unproven. Add the robots.txt rule, then check your logs to see whether it is honored - and if the block genuinely matters to you, plan on server- or firewall-level enforcement rather than trusting an undocumented crawler to read your directives.
The strategic question is narrower than for the Western training crawlers. Allowing PanguBot mostly affects how Huawei's models represent your domain to users inside China and its enterprise deployments, an audience that either matters to you a great deal or not at all.
How PanguBot behaves
- No operator documentation page, unlike Huawei's own PetalBot.
- Community registries record robots.txt compliance as unclear rather than confirmed.
- No IP feed or reverse-DNS convention, so traffic cannot be authenticated.
- Feeds PanGu multimodal models used across Huawei cloud and device products.
Allow PanguBot
Your content can be represented in Huawei's PanGu models and the cloud and device AI features built on them.
User-agent: PanguBot Allow: /
Block PanguBot
Keep your content out of Huawei's training data - and given the thin documentation, consider enforcing at the server rather than trusting robots.txt.
User-agent: PanguBot Disallow: /
How to verify PanguBot
No published verification method
Huawei publishes no IP range feed or reverse-DNS convention for PanguBot - a notable gap, since its own PetalBot has a documented verification domain. Do not assume the aspiegel.com convention carries over; there is nothing stating that it does. With no way to authenticate the user-agent, edge rules have to key on the string and on crawl behavior.
Check an IP against this botCommonly confused with PanguBot
PanguBot FAQs
Does PanguBot respect robots.txt?
Unconfirmed. Community registries record its compliance as unclear, and Huawei publishes no crawler documentation for it. Add the rule, then verify in your logs rather than assuming it is honored.
Why is PanguBot less documented than PetalBot if both are Huawei?
Different teams and different purposes. PetalBot serves Petal Search and has a webmaster-facing portal because site owners are its constituency. PanguBot collects training data and has no equivalent public surface.
Can I verify PanguBot traffic using PetalBot's aspiegel.com domain?
No. Nothing documents that convention as applying to PanguBot, so do not assume it carries over. There is no published verification method for this token.
What is PanGu?
Huawei's family of multimodal large language models - its counterpart to GPT and Gemini - deployed across Huawei cloud services and devices, with significant use in Chinese enterprise and government settings.
Can PanguBot read your page right now?
Test any URL and see exactly what AI crawlers receive.