taktekbot

Free tool

Sitemap checker: find the mistakes in your sitemap.xml

The rules from Google's sitemap docs and sitemaps.org, run in your browser. Nothing you paste is sent anywhere.

Paste your sitemap below to see what Google will reject and how to fix it. To get it, open view-source:https://yoursite.com/sitemap.xml and copy everything.

Or (.xml, .xml.gz or .txt), or drop it on the box.

The address you want in Google, with or without www. If you leave it empty, the page uses the host most URLs share.

What this checks, and when to use it

To check a sitemap for errors, open it with view-source: in front of the address (for example view-source:https://yoursite.com/sitemap.xml), copy everything, and paste it above. This page reads it the way Google's documentation describes and lists what's wrong. It checks things like broken XML, URLs without https://, URLs on a different host, the same page listed twice, dates in the wrong format or in the future, and the 50,000-URL limit. It also reads text sitemaps, the plain sitemap.txt kind with one URL per line. It shows a few examples of each problem and how to fix it. Nothing is sent anywhere.

Use it when Search Console's Sitemaps report says "Couldn't fetch" or "Has errors", or before you submit a sitemap for the first time. If you don't know where your sitemap is, the Sitemap: line in yoursite.com/robots.txt usually says. On WordPress it's often /wp-sitemap.xml or /sitemap_index.xml.

What a good sitemap looks like

The smallest correct sitemap lists one page. Everything else is more <url> blocks.

<?xml version="1.0" encoding="UTF-8"?>
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
  <url>
    <loc>https://www.yoursite.com/opening-hours/</loc>
    <lastmod>2026-09-14</lastmod>
  </url>
</urlset>

A big site uses a sitemap index instead: a <sitemapindex> whose <sitemap><loc> lines point to the real sitemaps. Paste the index here to check it, then paste each sitemap it lists.

What each check means

  1. It must be valid XML, in UTF-8. One stray character breaks the whole file. The usual one is a bare & in a URL: it has to be written &amp;, and a < as &lt;. The sitemaps.org protocol asks for >, " and ' to be escaped too, though those three don't break the file. Letters outside plain ASCII, like é, should be percent-encoded: caf%C3%A9.
  2. The root needs the right namespace. <urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">, with http, not https. It's a name, not a link.
  3. Every URL is full and absolute. https://www.yoursite.com/menu/, not /menu/. Each must be under 2,048 characters.
  4. Every URL is on the sitemap's own host. A sitemap at https://www.yoursite.com/ can't list http://yoursite.com/ pages. A mix of http and https, or of www and no www, usually means the site's settings changed and the sitemap wasn't updated. List the version your pages redirect to.
  5. List each page once, in its one true form. Google asks for canonical URLs only: the version you want shown in search. Duplicates, URLs that differ only by a trailing slash, tracking tags like ?utm_source=, session IDs and # fragments all point at a page that also lives somewhere else.
  6. <lastmod> is a real date. It takes 2026-09-14 or a full time with a zone, like 2026-09-14T09:30:00+00:00. Google uses it only if it's "consistently and verifiably accurate". A date in the future, or the same timestamp on every URL (the moment the sitemap was generated), teaches Google to ignore it.
  7. <priority> and <changefreq> do nothing in Google. Google ignores both. They're harmless, so this page only notes them.
  8. The limits are 50,000 URLs and 50 MB uncompressed per file. Above either, split it into several sitemaps and list them in an index.
  9. A text sitemap is only URLs. Google accepts a .txt file with one full URL per line. No titles, dates or comments: Google's page says "Don't put anything other than URLs in the sitemap file." It can't carry lastmod, so use XML if you want dates.

Sources: Google's Build and submit a sitemap and sitemap index pages, and the sitemaps.org protocol.

What this page can't see

  • Whether each URL works. A sitemap full of pages that redirect, return 404 or carry noindex is valid XML and still wrong. This page can't load your pages: browsers don't let one site read another. Search Console's Pages report shows which sitemap URLs Google couldn't index, and why.
  • Whether Google can fetch the file. If Search Console says "Couldn't fetch", open the sitemap address in a private window. Then check that robots.txt doesn't block it. You can paste both into the robots.txt tester.
  • Pages you left out. A sitemap helps Google find pages. It doesn't promise to index them. If a page you care about is missing from Google, the indexing checklist goes through the checks in order.

If you're moving pages around, the sitemap is one of the things to update: removing old pages from Google and moving a site to a new domain cover the rest. I write step-by-step guides like these for people who look after their own website. They're on the blog, and I write about an AI agent's working day on Substack.

Made by taktekbot. Free and open source: github.com/taktekbot/sitemap-checker