The most comprehensive crawl checker

Can AI crawl your site?

The most comprehensive AI crawl checker on the market. Sona tests 45 AI crawlers, GPTBot, ClaudeBot, PerplexityBot and 42 more, then catches what status codes hide: challenge pages, cloaked content, WAF, rate-limit, render and noindex blocks, with the exact fix for your CDN. Results in seconds.

https://

No credit card · Free forever · Every bot tested live

45 AI crawlersLive user-agent proberobots.txt per botWAF / CDN blocksCloaking & soft blocksllms.txt + sitemapCopy-paste fix
AI access reportyoursite.com
41/45bots
41 of 45 AI bots can reach you

4 crawlers are blocked. Fix the flagged rows to open full access.

GPTBotAllowed
ClaudeBotAllowed
PerplexityBotrobots.txt disallow
GooglebotAllowed
BytespiderWAF 403 block
Every blocked bot ships with a copy-paste robots.txt fix.
15%–25%

Estimated drop in organic web traffic from zero-click AI search

Source: Bain & Company, 2025

ChatGPT visitors convert better than traditional Google search

Source: Webflow, 2025

50%

Of U.S. consumers now intentionally use AI-powered search

Source: McKinsey AI Discovery Survey, 2025

94%

Of B2B buyers now use AI while researching a purchase

Source: Forrester, Buyers’ Journey Survey 2025

The problem

AI bots get blocked, and nobody tells you.

  • ×A single robots.txt disallow rule, or a broad wildcard, quietly locks out named AI crawlers
  • ×Your CDN or firewall returns a 403 to bot user-agents while browsers sail through
  • ×Bots get an HTTP 200 that is really a challenge page, a block-page redirect, or a stripped, cloaked response
  • ×Rate limiting throttles crawler traffic with a 429, even when robots.txt says welcome
  • ×JavaScript-only pages render blank to bots that do not execute scripts
The solution

Sona shows you which bot, and why.

  • Test all 45 AI crawlers, live probes with real published user-agents wherever the bot sends traffic
  • Parse your robots.txt per bot, wildcards included, and name the exact rule that blocks each one
  • Compare each bot against a normal browser to catch WAF, CDN, rate-limit, and cloaked soft blocks
  • Fingerprint the exact WAF/CDN vendor blocking bots and hand you the paste-ready allowlist fix
  • Hand you a corrected robots.txt block covering exactly the bots you are missing
The basics

What is AI crawlability?

AI crawlability is whether AI bots, GPTBot for ChatGPT, ClaudeBot for Claude, PerplexityBot, Googlebot and dozens more, can actually reach and load your pages. It is the first gate: if a crawler is blocked by your robots.txt, your firewall, or a rate limit, it never sees your content, no matter how good that content is.

Access is step one, not the finish line. Once AI bots can reach you, whether you get cited depends on your content and structure, which the deeper Sona AI Visibility audit measures. This free checker makes sure the door is open first.

How it works

Every bot. One question: can they reach you?

Paste your URL and Sona probes every crawler for you.

1

Enter your URL

Paste any webpage URL. One field, no credit card. Sona resolves the address and reads your robots.txt before probing.

2

Every bot tested live

We request your page with the real user-agent of every major AI crawler, GPTBot, ClaudeBot, PerplexityBot, Googlebot and 41 more, compare what each bot receives against a normal browser to catch challenge pages and cloaked responses, and evaluate your robots.txt bot by bot.

3

Get your access report

An access score, a pass/fail row per bot, the exact reason each blocked bot fails, a paste-ready fix for your specific WAF or CDN, per-section testing from your sitemap, and a prioritized action plan you can email or share with your team.

What we test

Access isn’t one check. It’s every way a bot gets blocked.

One tool, one job: access. Each check maps to a specific reason AI crawlers get turned away, and tells you exactly what to change.

Roster reviewed August 2026 · 45 bots and counting

Full Bot Matrix & Access Score

A pass/fail row for all 45 crawlers, each tested live, rolled into one headline access score.

robots.txt Diagnosis

Parsed per bot, wildcards included, naming the exact rule that blocks each crawler.

WAF Fingerprint + Paste-Ready Fix

Names the exact vendor blocking bots, Cloudflare, Akamai, DataDome, Imperva and more, and generates the allowlist rule for it.

Cloaking & Soft-Block Detection

A 200 is not a pass. We read what each bot receives and catch challenge pages, block-page redirects and cloaked responses.

Rate-Limit & Pay-Per-Crawl Detection

Spots 429 throttling that drops crawlers, and flags HTTP 402 pay-per-crawl walls as monetization, not misconfiguration.

Render-Block & noindex Check

Flags JS-only pages that render blank to bots, plus meta robots and X-Robots-Tag rules that turn crawlers away.

Whole-Site Section Testing

Samples one URL per path template from your sitemap, so a homepage that passes cannot hide a /docs that 403s GPTBot.

AI Policy Signals

Content-Signal, RSL licensing, noai, ai.txt, TDM-Rep and the nosnippet controls that gate Google AI Overviews.

Agent Readiness

Markdown content negotiation, llms-full.txt, MCP discovery and agents.json, the emerging agent conventions, checked.

llms.txt & Sitemap Health

llms.txt presence, spec conformance and link checks, plus broken sitemaps, 5xx robots.txt and malformed rules, with a free generator.

Bot IP Spoof Verifier

Up to 18% of bot traffic is fake. Verify any visitor IP against operator-published ranges and reverse DNS

Categories, Action Plan & Report

Bots grouped by training, retrieval, agent and stealth tier, ranked into a fix list, with a permanent shareable report link.

Real impact

See the difference the fix makes

Most sites quietly block a slice of the AI crawler roster. Here is a typical result, before and after the copy-paste fix.

Before
34/45
×PerplexityBot blocked in robots.txt
×GPTBot rejected by WAF with a 403
×Bytespider hitting a 429 rate limit
×Google-Extended disallowed by a wildcard
+11 bots
After the fix
45/45
PerplexityBot allowed in robots.txt
GPTBot user-agent whitelisted at the CDN
Crawler rate limits raised
Wildcard narrowed, all control tokens allowed
Reference

Know every bot before you test

Who runs each AI crawler, what it powers, its exact user-agent, and the robots.txt rules to allow or block it. A plain-English reference for 45 bots and control tokens, grouped by what each bot actually does, plus a stealth tier for agents that hide.

Stealth / undocumentedNo honored user-agent
xAI / GrokDeepSeekOpenAI OperatorDevinManus

These operators and agents publish no honored crawler user-agent, so no tool can pass or fail them by UA. Sona flags them as a distinct tier instead of giving false comfort.

Free tool suite

Access is one piece. Here’s the rest.

Each free Sona tool owns one job. Start by opening access, then move to the tool that fits your next question.

From snapshot to system

This check is a snapshot. Sona is the system.

This tool tells you which AI bots are allowed to reach your site right now. Sona Agent Analytics tells you which ones actually do, then follows the trail: which pages each crawler reads, how many human visitors each engine refers back, and which of those referrals turn into pipeline and closed-won revenue.

Covering every major answer engine
ChatGPTGeminiGoogle AI ModeAzure (OpenAI)PerplexityGrokAnthropicDeepSeek

Agent Analytics

“Which AI crawlers visit my site, and how do they drive revenue?”

See AI crawler activity by bot, hour and page, with verified bot signatures and anomaly alerts when a crawler appears or gets blocked
Track referrals per page: how many human visitors each AI engine sends after crawling it, broken down by OpenAI, Anthropic, Perplexity and more
Connect crawl activity to pipeline and closed-won revenue, so the pages AI reads most are prioritized by the revenue they drive

AI Search Insights

“Once they can reach me, am I actually getting cited in answers?”

A daily visibility score and share of voice across ChatGPT, Gemini, Google AI Mode, Azure (OpenAI), Perplexity, Grok, Anthropic and DeepSeek
Sentiment scoring and the citations driving each answer, so you know which sources to invest in
Every competitor ranked on every prompt, with alerts when a watched rival moves
Get crawledGet readGet citedRevenue
FAQ

Common questions about AI crawlability

Everything GTM and web teams ask before running their first crawl check.

An AI Crawl Checker tests whether AI crawlers, GPTBot (ChatGPT), ClaudeBot (Claude), PerplexityBot (Perplexity), Googlebot, CCBot and a dozen more, can reach your website. It requests your page live with each bot’s real user-agent, evaluates your robots.txt per bot, and returns a pass/fail matrix with the exact reason each blocked bot fails. If AI bots cannot reach you, you cannot show up in AI answers.

Breadth versus depth. The AI Crawl Checker is wide, shallow and instant: it tests every major AI bot against one URL and answers a single question, which bots can reach you and what is blocking the ones that cannot. The AI Visibility audit is one bot but deep: 25+ checks across crawlability, content structure, content quality, performance, security and accessibility. Run this tool first to make sure access is open, then run the full audit.

The full matrix is 45 bots: GPTBot, OAI-SearchBot, OAI-AdsBot and ChatGPT-User (OpenAI); ClaudeBot, Claude-SearchBot, Claude-User and Claude-Code (Anthropic); PerplexityBot and Perplexity-User (Perplexity); Googlebot, GoogleOther, Google-Extended, Google-CloudVertexBot, Google-NotebookLM, Google-Agent, Gemini-Deep-Research and GoogleAgent-URLContext (Google); Bingbot (Microsoft); Applebot and Applebot-Extended (Apple); CCBot (Common Crawl); Bytespider and TikTokSpider (ByteDance); Amazonbot, Amzn-SearchBot, Amzn-User and bedrockbot (Amazon); DuckAssistBot (DuckDuckGo); Meta-ExternalAgent, Meta-ExternalFetcher, Meta-WebIndexer and FacebookBot (Meta); MistralAI-User (Mistral); AI2Bot (Allen Institute); PetalBot and PanguBot (Huawei); cohere-training-data-crawler (Cohere); Diffbot; YouBot (You.com); ExaBot (Exa); Kimi-User (Moonshot AI); TongyiBot (Alibaba); ChatGLM-Spider (Zhipu AI); and Webzio-Extended (Webz.io). Bots that send real traffic are probed live with their published user-agents; control tokens like Google-Extended are evaluated against your robots.txt; and fetchers documented to ignore robots.txt are judged by live probe only. Operators with no honored user-agent, like xAI’s Grok, DeepSeek, OpenAI’s Operator, Devin, Manus and Tavily, are flagged as a separate stealth tier rather than given a misleading pass/fail row.

It is a simple headline: how many of the tested AI bots can reach your page, for example "41 of 45 AI bots can reach you." A bot counts as blocked if your robots.txt disallows it, or if a live request with its user-agent is rejected (403 or 429) while a normal browser gets through. For every blocked bot, the report shows the specific failure mode and the fix.

Seven failure modes cover almost every case: 1) a robots.txt disallow rule for the bot (or a wildcard rule that catches it), 2) a WAF/CDN block that rejects the bot's user-agent, 3) a soft block, the bot gets HTTP 200 but the body is a challenge page or a login/block-page redirect, 4) cloaking, the bot receives a stripped page with a fraction of the content a browser gets, 5) rate-limiting that throttles crawler traffic, 6) render-blocking, the content only exists after JavaScript runs, which AI crawlers don't execute, and 7) a noindex directive in a meta tag or X-Robots-Tag header. There's also pay-per-crawl (HTTP 402), which the report labels as deliberate monetization rather than a failure. The report identifies which one applies to each blocked bot and gives the targeted fix.

Because modern bot protection rarely sends an honest 403. Cloudflare, DataDome, and similar services often return HTTP 200 with a challenge page ("Just a moment...") that contains none of your content, or serve bots a stripped version of the page. Status-code-only checkers mark these as passing. This tool reads the body each bot actually receives and compares it against what a normal browser gets on the same URL: challenge-page signatures, block-page redirects, and content that shrinks to under a quarter of the browser version all get flagged, with the vendor-specific fix.

Yes. Beyond per-bot access, every report checks the discovery files crawlers read before your pages: robots.txt health (5xx responses, malformed rules, crawl-delay traps), sitemap presence and reachability, and llms.txt, whether it exists, conforms to the spec, and whether its links resolve. It also samples one URL per path template from your sitemap and re-tests representative bots against each, so a /blog that passes doesn't hide a /docs that blocks. If you do not have an llms.txt yet, the free llms.txt generator crawls your site and produces an editable, validated file ready to deploy.

Two newer layers the report covers beyond raw access. Policy signals are the emerging standards that declare how AI systems may use your content: Content-Signal directives in robots.txt, RSL licensing (the License directive backed by Reddit, Yahoo, and 50+ publishers), noai directives, ai.txt, TDM-Rep, and the nosnippet/max-snippet preview controls that quietly remove pages from Google's AI Overviews. Agent readiness covers the conventions that make your site cheaper for AI agents to consume: serving markdown via content negotiation (Accept: text/markdown), llms-full.txt, an MCP discovery card at /.well-known/mcp.json, and agents.json. None are required, the report shows where you stand on each.

Firewalls and CDNs. Services like Cloudflare ship bot-protection settings that return 403 or 429 to AI crawler user-agents regardless of what robots.txt says. The checker catches this by comparing each bot’s response against a normal browser request: if the browser succeeds and the bot fails, the block lives in your WAF or CDN configuration, not in robots.txt. The report fingerprints the exact vendor doing the blocking, Cloudflare, Akamai, DataDome, Imperva, HUMAN (PerimeterX), AWS WAF, Sucuri and others, and generates a paste-ready allowlist rule for that specific vendor.

Yes, use the free Bot IP Verifier at /verify-bot. User-agent strings are trivially faked (studies find 12-18% of "Googlebot" traffic is spoofed), so the verifier checks a visitor IP against the bot operator's published IP range feeds (OpenAI, Google, Microsoft, Perplexity) or documented reverse-DNS hostnames with forward confirmation (Apple, Amazon, DuckDuckGo, Huawei). You get a clear verdict: genuine, spoofed, or, for operators like Anthropic that publish no verification data, honestly unverifiable.

Yes. Enter a URL and you get the full report: access score, the complete bot matrix, per-bot failure diagnosis, a copy-paste robots.txt fix, and a prioritized action plan you can have emailed to you or share with your team via a permanent report link. For continuous monitoring of which AI bots actually visit your site, check out Sona Agent Analytics.

Agent Analytics is part of the paid Sona platform. This free tool tells you which AI bots are allowed to reach your site; Agent Analytics tells you which ones actually do, and what that traffic is worth. It logs every AI crawler visit by bot, page and hour with verified bot signatures that separate real crawlers from spoofed user-agents, then connects that activity to outcomes: how many human visitors each AI engine refers back after crawling you, and which of those referrals turn into pipeline and closed-won revenue. That lets you prioritize the pages AI reads most by the revenue they drive. It is the difference between permission and reality.

Want to go beyond one-time checks?

Your free access report is step one. Monitor which AI crawlers actually visit your site, and how every crawl turns into new opportunity, with Sona Agent Analytics.