Skip to content
AuditAE / AI Crawler CheckerFree · No login

Free AI crawler checker: is your site
blocking AI bots?

Enter a domain. We read your robots.txt for 33 bots, including the 14 AI and search bots that decide the result, check your homepage’s meta robots tag and X-Robots-Tag header, and test whether a firewall or CDN turns away GPTBot, ClaudeBot, PerplexityBot and ChatGPT-User. Every verdict quotes the line that decided it, and every problem comes with a paste-ready fix.

Test a specific path (optional)
33 bots · robots.txt, page tags and firewall · no card, no email
Bots we check

33 bots, and only 14 of them decide the headline.

The headline counts the 14 AI and search bots that fetch pages for AI search and for people’s questions (8 AI search bots, 6 user-triggered fetchers). Blocking one of the 8 training bots is a policy choice, so it’s shown as a neutral count. Bingbot is a search crawler rather than an AI bot; it’s here because Microsoft Copilot sends its web searches to Bing. Each name links to what the bot does in our guide to AI crawlers.

AI search bots (8)

User-triggered fetchers (6)

Training bots (policy choice) (8)

Other bots, not counted (11)

Purposes are our classification, checked against each vendor’s own crawler documentation. Bots we couldn’t source, legacy tokens and brand monitors sit under “Other bots, not counted”: you see their verdict, but they never change the result.

Allowed in robots.txt, still blocked

robots.txt is a request. A firewall rule is enforcement.

A CDN or web application firewall sits in front of your site and can block or challenge a request by its user-agent before robots.txt ever comes into it. A bot rule, a “block AI bots” switch or an aggressive challenge setting can turn AI crawlers away while robots.txt says they’re welcome.

That’s why the check loads your homepage once as a normal request and then as GPTBot, ClaudeBot, PerplexityBot and ChatGPT-User. If the bot requests get a 401, 403, 429, 503 or a challenge page and the normal one doesn’t, we flag a likely firewall or CDN block. We can send each bot’s user-agent but not its vendor’s IP addresses, and some CDNs let the real bots through while stopping look-alikes, so treat it as a strong hint and confirm in your CDN’s bot settings.

Our probe requests are labelled, so they’re easy to spot in your logs. About our bot.

Then the other half

Access is not citation.

A bot that can reach your pages can cite them. Whether ChatGPT, Perplexity, Gemini or Google AI Overviews actually do depends on your content, the question and the sources each engine trusts. The free citation check runs real buyer prompts through the engines and shows which of them cite you and which cite someone else.

On WordPress? The free AuditAE plugin shows which AI crawlers actually visit, audits robots.txt inside wp-admin and can allow AI bots in one click. For robots.txt itself, see the WordPress robots.txt guide.

AI crawler checker questions

What each bot is for is covered in AI crawlers explained; these answers are about the checker.

Your robots.txt names the bot by part of its name, for example “User-agent: claude” rather than “User-agent: ClaudeBot”. Crawlers that match partial names would apply that group; crawlers that follow the robots.txt standard (RFC 9309) match the full name and fall back to the “User-agent: *” group. When those two readings give different answers, we show Unclear with both lines instead of guessing. Name the bot in full to remove the doubt. For the bots it scores, our AI readiness audit counts a partial-name block as blocked, to be safe.

The bot can reach your site but is disallowed from some paths, such as /wp-admin/ or /private/. That's usually deliberate. Use the path test under the form to check one page for each AI and search bot.

Robots.txt is a request; a CDN or firewall rule is enforcement. We load your homepage once as a normal request and four more times with the user-agents of GPTBot, ClaudeBot, PerplexityBot and ChatGPT-User. If those get a 401, 403, 429, 503 or a challenge page while the normal request didn't, something between the bots and your site is stopping them. We send the user-agent but not the vendor's IP addresses, and some CDNs let the real bots through while stopping look-alikes, so the result is a strong hint, not proof. Check your CDN's bot settings to confirm.

We follow RFC 9309, the robots.txt standard. A 404, 410 or other 4xx means there are no rules, so every bot is allowed. A 5xx server error means crawlers treat the whole site as blocked until it recovers, so we show that in red. A 401, 403 or 429 is an answer to our request (HTTP auth, a firewall or a bot rule), so we can't tell what AI crawlers see and say so. A web page served at /robots.txt counts as no rules.

Blocking a training crawler such as GPTBot or ClaudeBot, or the Google-Extended token, is a policy choice about model training, not a visibility problem, so it never turns the result red. The headline counts only the 14 AI and search bots that fetch pages for AI search and for people's questions. Training blocks are shown as a neutral count.

Use the “Block training, keep AI search access” fix in the result: it disallows the training bots and allows the search and user-triggered ones, keeping your existing path rules. The cost: training data is also how AI learns your brand, so blocking it can reduce mentions of you that appear without a link. Run the free citation check to see whether AI mentions you before you decide.

Check whether AI mentions you

Not necessarily. Access is the precondition, not the result. Whether ChatGPT, Perplexity, Gemini or Google AI Overviews cite you depends on your content, what else answers the question and which sources each engine trusts. The free citation check runs real prompts through the engines and shows who they cite.

Run the free citation check

We cache the result for up to 10 minutes (1 minute if the firewall check was skipped) so a re-check is instant, and log the domain checked and summary counts. The domain is also kept for an hour as a rate-limit key, so no single site gets checked too often, and your IP address is kept for 24 hours for the per-visitor limit. We don't store the page content.

Three OpenAI user-agents with three jobs: training, search, and fetching a page a person asked about. The crawler guide explains each one.

AI crawlers explained

OpenAI documents GPTBot as its training crawler and OAI-SearchBot as the one that surfaces sites in ChatGPT search, so they're separate decisions. The crawler guide covers the details.

AI crawlers explained

Free citation check

Can AI reach you? Now see whether it cites you.

Four engines, your brand, one free check. See who ChatGPT, Perplexity, Gemini and AI Overviews cite for your buyers' questions.

Run a free auditFree
Or preview one first.

We check all 4 engines

1 prompt × 4 engines · ~30 seconds · no card required

Reachable is step one. Cited is the goal.

Sign up free to run citation audits across ChatGPT, Perplexity, Gemini and Google AI Overviews, with a fix-it playbook after each one.