Sitemap Audit

Find the Errors Google Reports, and the Pages That Cannot Be Indexed

Audit a Sitemap

Fetch

Try one:

tools.scoreroute.com
This site
sitemaps.org
The protocol
w3.org
Standards body
example.com
No sitemap

Finding the sitemap...

Nothing is stored. The files are read, checked and discarded.

Why a Valid Sitemap Can Still Index Nothing

Ask a chatbot why your pages are not indexed and it will give you a confident answer about sitemap quality. The problem is that a valid sitemap says nothing about whether a single page is indexable. Validating the XML is like checking that a letter is addressed correctly. It says nothing about whether the door is locked, whether the recipient moved house last year, or whether the postman was told to skip that street.

Those are the real failures, and every one of them is invisible to a file validator. A page can return 200, sit inside a perfectly formed urlset, and still be dropped from the index because its head carries a meta robots noindex. It can be dropped because a rel=canonical tag points somewhere else, in which case Google keeps the other URL and files yours as an alternate. It can be dropped because robots.txt disallows that path, which means the sitemap is asking for something the crawl rules forbid. It can be dropped because it redirects, because it 404s, or because the server sends an X-Robots-Tag header that says noindex before the HTML is even read.

This tool fetches the pages, not just the file. It takes the sitemap, spreads a sample evenly across it, and requests each URL without following redirects so a 301 stays visible as a 301. For every page it reports the status, the header, the meta tag and the canonical target, and pairs each verdict with the label Search Console uses for that class of problem so you can match it against the report you already have.

What You Get

  • Sitemap discovery — reads the Sitemap: line in robots.txt first, and falls back to probing six common paths. One redirect hop is followed, so a sitemap that only exists on the www host is still found instead of being reported as missing.
  • The structural rules Google actually enforces — 50MB uncompressed, 50,000 URLs, fully-qualified absolute loc values, the W3C datetime formats for lastmod, the required namespace, leading whitespace, duplicate tags, nested index files, and URLs that sit on a different site than the file they are listed in.
  • Per-page indexability — status code without following redirects, X-Robots-Tag, meta robots noindex, a canonical pointing at another URL, and a robots.txt Disallow resolved with the same RFC 9309 precedence a crawler uses, so a path cannot be reported as blocked in one place and allowed in another.
  • Duplicates that do not look like duplicates — a trailing slash, a www prefix or a stray utm parameter makes two entries the same page to a crawler. The audit names the pair so you can fix it in the generator rather than in the file.
  • A cleaned sitemap.xml — the audited file with every URL proven to be dead removed, ready to download.
  • JSON output — the full report as a copyable object, for piping into your own tooling or an agent.

What This Tool Will Not Tell You

Google is direct about what a sitemap is worth. Submitting one is a hint. It does not guarantee Google downloads it, and a URL listed in a sitemap is not guaranteed to be crawled or indexed. A sitemap makes discovery faster for pages that are already eligible. It does not make an ineligible page eligible.

The per-page results are also a sample, and the ceiling is a hard one. Free Cloudflare Pages Functions allow 50 outbound requests per request, and the run already spends some of them finding the file. The sample is spread evenly across the whole sitemap rather than taken from the front, because the first twenty entries on a real site are the homepage and its immediate children, and a front-of-list sample reports a clean bill of health on a site that is broken further in. Every page that was not fetched is listed without a live verdict. It is never marked as fine.

Frequently Asked

My sitemap is valid. Why is nothing indexed?

Because a valid sitemap only means the file parses. Google still decides per page whether to index it. The usual causes are a meta robots noindex on the page, a canonical tag pointing at a different URL, a Disallow rule in robots.txt covering that path, a redirect, or a 404. All five return HTTP 200 or 3xx and are invisible to a validator that only reads the XML. This tool fetches the pages, not just the file, and reports which of these is happening to which URL.

Does a sitemap get my pages indexed?

No, and Google is explicit about it. Submitting a sitemap is a hint. It does not guarantee Google downloads it, and a URL listed in a sitemap is not guaranteed to be crawled or indexed. What a sitemap reliably does is speed up discovery of pages that are already indexable. If a page is blocked, redirected or marked noindex, adding it to the sitemap changes nothing.

Does lastmod in a sitemap help my rankings?

Not directly. Google states it uses the lastmod value when it is consistently and verifiably accurate, for example when compared against the actual modification time of the file. That condition is the whole catch. One identical timestamp on every URL in a large sitemap is the signature of a build time, not a modification date, and it teaches the crawler nothing. Set lastmod per page to the date that page last changed, or drop the tag rather than fill it with today's date.

Should I keep priority and changefreq in my sitemap?

They are valid, and Google states plainly that it ignores both values. They add bytes to every entry and buy nothing. The common belief that priority 1.0 gets a page crawled first is not how it works. Drop them and put the effort into lastmod, which is at least read.

How many URLs should one sitemap hold?

A single sitemap by Google is limited to 50,000 URLs and 50MB uncompressed data. Beyond that, it needs to be segmented, with the components documented in a sitemap index file. An index file can contain 50,000 entries, must have sitemaps referenced in its own directory or less, and is not authorized to list other index files, including itself. Google limits users to a maximum of 500 index files per site.

Why are two of my URLs flagged as duplicates when they look different?

A crawler treats a trailing slash, a www prefix and tracking parameters as the same page. /product and /product/ are one URL, and so are example.com/x?utm_source=a and example.com/x?utm_source=b. Listing them twice splits your signals, Google picks the version it prefers, and the other is reported as a duplicate without a user-selected canonical. Fix it at the source: one URL in the sitemap, and a rel=canonical on the page pointing at the version you want.

How many URLs does this actually check?

The whole file is parsed, so every URL is checked against the structural rules. A sample of the pages is then fetched live, spread evenly across the file rather than taken from the front, because on a real site the first entries are the homepage and its immediate children. The sample size is on the control above, up to 35. That ceiling is a limit of free Cloudflare Pages Functions, which allow 50 outbound requests per request. Every page not fetched is listed without a live verdict rather than assumed fine.

Is the domain I check stored or logged?

No. The request is answered with a no-store cache header, the files are read and discarded, and nothing is written to storage or to disk. We do not keep a list of the sites people audit.

Comments & Ratings

Be the first to comment.

References