AI Crawler Log Analyzer

Drop In the Access Log, Find Out What GPTBot and ClaudeBot Actually Did

Load Your Access Log

Drop a log file here, or click to choose one
.log · .txt · .gz · .json  —  Apache, nginx, IIS W3C, Cloudflare Logpush

Ask a Chatbot Which AI Crawlers Hit Your Site

A chatbot will tell you. It will also be wrong. It cannot see your access log. It has read a lot about GPTBot and nothing about the 41 lines in last night's nginx log where ClaudeBot got a 403 on every request. A language model cannot count rows, group them by user agent, or tell a real crawler from a scraper wearing its name. What it can do is describe the grep command you should have run.

This tool runs that command properly. Drop in the log and you get the numbers a pager would have given you: which AI crawlers arrived, how often, what they asked for, what status each one got, and whether the requests actually came from the operator they claim. Training crawlers and retrieval crawlers are counted separately, because they are separate decisions. A GPTBot hit is index input. A Perplexity-User hit is a citation.

What You Get

  • Five log formats, detected for you — Apache common, Apache combined, nginx, IIS W3C with its Fields header, and Cloudflare Logpush in NDJSON. Gzipped logs are unpacked in the browser with the platform decompressor, so access.log.1.gz works the same as access.log.
  • A per-crawler table — requests, unique URLs, bytes, first and last sighting, and a 2xx/3xx/4xx/5xx breakdown for each of 55 named agents.
  • A crawler-by-page matrix — the honest answer to "what do these engines actually read on my site", which no summary of totals can give you.
  • Reverse DNS spoof checks — every distinct IP behind an AI crawler can be looked up, and the result is a verdict: the address points back at the operator's own domain, or it does not. Between five and eight per cent of requests claiming to be a known AI crawler are spoofed, and a block list built on the user agent string alone is blocking the wrong traffic and reading the wrong numbers.
  • JSON output — the whole report as a file, for a ticket, a dashboard or an agent.

What This Will Not Do

It does not crawl your site and it does not fetch anything you did not give it. It cannot tell you whether the pages a crawler read are the pages you wanted read, or whether what it found there was worth citing. It reads one file and counts. The robots.txt side of the same question — what you meant rather than what happened — is a different tool.

Frequently Asked

Why is my access log the only place this answer exists?

Because AI crawlers do not run JavaScript. Google Analytics counts referrals, which are the people who clicked a citation, and never sees the crawl itself. Search Console has no field for GPTBot or ClaudeBot at all. The request is a row in a file on your own disk, and that file is the complete record.

What is the difference between GPTBot and OAI-SearchBot?

GPTBot is OpenAI's training crawler. OAI-SearchBot and ChatGPT-User are retrieval agents, called when someone asks ChatGPT a question or pastes your link. A GPTBot fetch is index input. An OAI-SearchBot fetch is a lookup. If you block GPTBot and leave OAI-SearchBot open you stay out of training and keep your citations. If you do the reverse you feed the model and lose the answer.

Why does Google-Extended never appear in my log?

Because Google-Extended is not a crawler. It is a token in robots.txt that switches Gemini grounding and training on or off. Google fetches with Googlebot and obeys the token itself, so no separate user agent is ever sent. If you go looking for a Google-Extended line in a log you will never find one, and you may conclude Google is ignoring you when it is not.

How does the spoof check work?

Every user agent is a claim anyone can make. For each distinct IP behind an AI crawler the tool asks a public DNS-over-HTTPS resolver for the PTR record, that is the name the address points back to, and compares it with the domain the operator documents. A match is a verdict of verified. A mismatch, or no record at all, is unverified. It is a signal and not proof, because a PTR record can be spoofed too, so the raw name is always shown next to the verdict.

Is my log uploaded anywhere?

No. The file is read with the browser file reader and parsed on the page. Nothing is sent to us and nothing is written to disk. The one exception is the reverse DNS button, which you press deliberately: it sends only the crawler's public IP address to a public DNS resolver, never your log, never the URLs in it and never your own traffic.

Which log formats does it read?

Apache common, Apache combined, nginx in either form, IIS W3C where the Fields header is present, and Cloudflare Logpush in NDJSON. Gzipped files are unpacked in the browser with the platform decompressor, so access.log.1.gz works the same as access.log. The format is detected from the first lines rather than asked for.

My lines did not parse. What now?

The report counts unparsed lines and prints the first three, which is usually the whole diagnosis. A log in true Apache common format has no user agent column, so no bot can be identified and every line lands in the unknown bucket. That is reported as a finding rather than left as an empty table.

Does this tell me whether my robots.txt is right?

No. This reads what happened. Your robots.txt says what you meant. Resolving a robots file against 19 AI crawlers with the right precedence is the job of the AI Crawler Access Checker, linked at the bottom of this page, and the two make a pair: one shows intent, this one shows what the server actually did.

Pairs Well With

  • AI Crawler Access Checker — resolves your robots.txt against 19 AI crawlers, so you know what you intended before you read what happened.
  • SEO Checker — reads one page the way a crawler would, including the structured data and the tags it is indexed by.
  • Sitemap Audit — finds the URLs a crawler cannot reach, which is the other half of any crawl gap.
  • User Agent Parser — identifies a single user agent string when you only have that and no log.

Comments & Ratings

Be the first to comment.

References