# AI Crawler Access Checker

> Who Can Read Your Site, and What Breaks llms.txt

Free AI crawler access checker. Audit robots.txt against 19 AI bots and validate llms.txt to spec. See the exact rule that blocks GPTBot. No signup.

URL: https://tools.scoreroute.com/tools/ai-crawler-checker/

Markdown: https://tools.scoreroute.com/tools/ai-crawler-checker/.md

## Audit a Domain

Fetching robots.txt and llms.txt...

Nothing is stored. The three files are read, checked and discarded.


### Why ChatGPT's Answer About Your Own Robots.txt Is Usually Wrong

Ask a chatbot whether GPTBot can read your site and it will give you a confident answer built from general knowledge. The problem is that the answer depends on three things no language model can look up for you: the exact contents of a file on your server, the group and path rules in it, and whether the second file your setup depends on is even reachable.

robots.txt precedence is the part people get wrong most often. A crawler only falls back to the `User-agent: *` group when it finds no group named after it. A named group beats the wildcard every time, whatever the wildcard says. Inside a group, the longest matching path wins, and when an Allow and a Disallow tie on length, Allow wins. A site with a blanket `Disallow: /` and a named GPTBot group is open to GPTBot. A site with `Disallow: /` and `Allow: /` in the same group is open to everything, because the Disallow does nothing.

Then there is the file the second half of your setup depends on. Publishing `/llms.txt` does nothing at all if the AI crawlers you care about are blocked from fetching it, and the most common way that happens is a `Disallow: /` in the wildcard group that nobody thought about because the file loads fine in a browser.

This tool fetches your three files and resolves the rules the way a compliant crawler would, then shows you the specific line responsible for each verdict.


### What You Get

- **Per-crawler verdicts for 19 documented AI bots** — GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-SearchBot, Claude-User, anthropic-ai, PerplexityBot, Perplexity-User, Googlebot, Google-Extended, Bingbot, Applebot-Extended, meta-externalagent, CCBot, Amazonbot, cohere-ai, Bytespider and MistralAI-User, each marked as search or training so the two are never confused again.
- **The deciding rule, not a yes or no** — every verdict carries the exact Allow or Disallow line that produced it, and whether it came from that crawler's own group or the wildcard.
- **llms.txt spec validation** — single H1, summary blockquote, sections that contain links, absolute URLs, no duplicates, no empty sections, and a check for the mistake that hides everything else: an HTML page served as llms.txt with a 200 status.
- **A contradiction report** — a published llms.txt that no search crawler can reach, a site that blocks OpenAI training but allows its search bot (and the reverse, which is the mistake), a blocked Googlebot, a missing llms-full.txt, or an H1 that does not match the domain.
- **Copy-paste fixes** — a ready robots.txt AI crawler group and an llms.txt starter built for your domain, both in a textarea you can copy straight out.


### What This Tool Will Not Tell You

llms.txt is a community convention from llmstxt.org. It is not a standard, and no engine has adopted it as a ranking input. Google has said plainly that it does not use llms.txt for indexing or ranking. Publishing one is cheap and it does give an agent a curated map of your site, but it is not a switch that makes you appear in ChatGPT, and any tool that tells you otherwise is selling you a promise it cannot keep.

What this audit does tell you is the part that is real today: whether the crawlers that already exist are permitted to read you, and whether the file you published can actually be reached by the agents you had in mind. Fixing an access problem is worthwhile on its own. Expecting a traffic jump from it is not.


### Frequently Asked

**Does llms.txt actually affect Google or ChatGPT rankings?**

No, and anyone selling you a tool that promises otherwise is selling you something. llms.txt is a community convention proposed on llmstxt.org, not a standard any engine has adopted as a ranking input. Google has stated that it does not use llms.txt for indexing or ranking. The file is worth publishing because it is cheap, it tells an agent which of your pages matter, and several smaller crawlers and agent frameworks do read it. It is not a switch you flip to appear in ChatGPT. What does have an effect today is robots.txt, because that is the one file every crawler already obeys, and it is where sites most often get the decision backwards by accident.

**What is the difference between GPTBot and OAI-SearchBot?**

They are separate crawlers with separate jobs and separate robots.txt user-agent tokens. GPTBot is the training crawler: it reads pages to build and improve models. OAI-SearchBot is the search crawler: it is what makes your pages eligible to appear in ChatGPT search results. ChatGPT-User is a third one, fetched when a user pastes a link into a chat. Blocking GPTBot while allowing OAI-SearchBot is the most common setup on the web and it is usually deliberate at large sites. The mistake is doing it the other way round, which allows training but silently removes you from answers.

**My robots.txt has Disallow: / in the wildcard group. Are AI bots blocked?**

Not necessarily, and this is where most advice goes wrong. A crawler first looks for a group whose user-agent token matches its own name exactly. Only if it finds no such group does it fall back to the wildcard group. So a site with a blanket Disallow: / plus a named group for GPTBot is open to GPTBot and closed to everything else. This tool resolves that per bot and shows you the specific line that decided it, rather than asking you to work it out by hand.

**Will blocking Google-Extended hurt my Google Search rankings?**

No. Googlebot controls Google Search, and Google-Extended is a separate token that covers Gemini training and grounding. Blocking Google-Extended leaves your Google Search and AI Overviews visibility untouched. The reverse is not true: if you block Googlebot you are out of Google entirely, and no amount of llms.txt work brings that back. If you see a site claiming Google-Extended is required for AI search results, that claim is wrong.

**What is the current Anthropic crawler token, ClaudeBot or anthropic-ai?**

ClaudeBot is the current token and it covers both training and retrieval. anthropic-ai was the earlier agent token and is still found in a lot of robots.txt files. A file that only names anthropic-ai is not naming ClaudeBot, so ClaudeBot falls through to the wildcard group and gets whatever that says. This tool lists both and reports which group each one is actually being judged by.

**Why does my llms.txt report an error when it clearly loads in my browser?**

Three usual causes. The URL returns your 404 page with a 200 status, so the file is really an HTML page and an agent reading it gets markup instead of content. The file was uploaded as JSON or YAML when the convention expects plain markdown. Or the file has no H1 line, which is the one part of the spec that is not optional. All three load fine in a browser and all three are invisible to an agent.

**Do I need llms-full.txt as well as llms.txt?**

It is optional, and the split exists for a reason. llms.txt is meant to be a short index a model can read cheaply. llms-full.txt holds the actual content for agents willing to spend the tokens. A one-line llms.txt with no full version gives an agent a title and nothing to act on, which is why this tool flags the missing file rather than passing it.

**Is the domain I type stored or logged anywhere?**

No. The request is answered from the Worker with a no-store cache header, the three files are read and discarded, and nothing is written to KV or to disk. We do not keep a list of the sites people audit.

**Why does my domain fail to audit?**

Three common reasons. The site only answers on one scheme and the other request timed out, in which case re-running usually gets through. The host blocks requests that do not look like a browser, and the file is never delivered. Or the site does not exist at that name, which is reported as unreachable rather than as a clean bill of health, so an unreachable domain never scores well by accident.


#### Related Tools

- [HTTP Status & Header Checker](/tools/http-status-checker/)
- [Website Down Checker](/tools/website-down-checker/)
- [SSL Certificate Checker](/tools/ssl-checker/)
- [URL Metadata Extractor](/tools/url-metadata-extractor/)

## References

- [IETF — RFC Index](https://www.rfc-editor.org/)
- [IANA — Protocol Registries](https://www.iana.org/protocols)

