taktekbot

Free tool

Noindex and canonical checker: is your page telling Google to skip it?

Google's robots meta tag and canonical rules, run in your browser. Nothing you paste is sent anywhere.

Paste a page's source below to find every tag or header that keeps it out of Google, and the ones search engines ignore.

Open the page in Chrome, Edge or Firefox, press Ctrl+U (⌘+⌥+U on a Mac, ⌘+U in Firefox), select all and copy. On a phone, type view-source: before the address in Chrome.

In a terminal: curl -sIL https://yoursite.com/page/. On Windows PowerShell: curl.exe -sIL https://yoursite.com/page/. The L follows redirects; paste everything and I read the last response. Or in the browser: open developer tools (F12), Network tab, reload, click the first row, and copy the Response Headers.

Where a page can say "leave me out"

A page can ask Google to leave it out of search in three places: a robots meta tag in its HTML, an X-Robots-Tag header sent with it, or a rel="canonical" link that names another URL as the main one. Paste the page's source above, and its headers if you have them, and this page lists every one of those it finds, what each one tells Google and Bing, and the ones that are written in a way the search engine ignores.

Use it when Search Console says a page is "Excluded by ‘noindex’ tag" or "Alternate page with proper canonical tag" and you can't see why, or when a page you published weeks ago still isn't in Google. Nothing you paste is sent anywhere.

What each tag does

  • <meta name="robots" content="noindex"> keeps the page out of every search engine that reads it. none means the same as noindex, nofollow. With name="googlebot" it's for Google only, with name="bingbot" for Bing only. Google also reads robots meta tags in the body, not just the head.
  • X-Robots-Tag: noindex is the same instruction, sent by the server as a header instead of written in the page. It's how PDFs and images get a noindex. You can't see it in the page source, which is why a forgotten one is hard to find: a server, CDN or plugin setting adds it to every response.
  • unavailable_after with a date that has passed works like noindex in Google. A date it can't read is ignored. It's a Google rule; SEO guides list it as Google-only.
  • <link rel="canonical" href="…"> says which URL is the main copy of this page. If it names another URL, Google usually shows that one and lists this one in Search Console as "Alternate page with proper canonical tag". That's right for a duplicate and a mistake for a page you want found. Google takes it as a strong hint, not an order, and only reads it in the <head>.
  • <meta http-equiv="refresh"> sends the visitor to another URL. Google treats a 0-second refresh as a permanent redirect and a delayed one as a temporary redirect, so it may index the other URL instead.

Mistakes this catches

  • A leftover noindex from a staging site, a "coming soon" mode, or a CMS's "discourage search engines" box.
  • A canonical Google never reads. The <head> ends as soon as the browser meets something that belongs in the body: an <img> tracking pixel outside <noscript>, an <iframe>, a <div>, or stray text. Everything after it, canonical included, lands in the body, and Google ignores a canonical there. This page parses your source the way a browser does and shows the tag that closed the head.
  • A canonical pointing at the wrong URL: every page pointing at the home page, http:// on an https:// site, the www twin, a staging host, a #fragment.
  • A relative canonical like /services/. Google accepts it but recommends the full address, because a copy of the site on another host would then claim to be the original.
  • More than one canonical, or a header and a tag that disagree. Keep exactly one.
  • noindex and a canonical to another page together. They ask for two different things. Google's advice is to use the canonical alone to pick the main copy, because noindex drops the page from search completely.
  • A noindex that JavaScript removes later. When Google sees noindex in the original source, it may skip running the page's JavaScript, so a script that removes it can come too late. If you want the page indexed, the noindex can't be in the source at all.

What this page can't see

  • robots.txt. If robots.txt blocks the page, Google never fetches it and never sees any of these tags, so a noindex there does nothing. Check the URL in the robots.txt tester.
  • What JavaScript adds. The source is what the server sends. If your site builds its pages with JavaScript, a tag can be added or changed after that. In Search Console, URL Inspection → Test live URL → View tested page shows the HTML Google ended up with; paste that here to check it too.
  • What your server sends to Googlebot. Some firewalls and plugins answer bots differently from people. URL Inspection's live test shows Google's view.

The rules here come from Google's robots meta tag and X-Robots-Tag specification, its guide to canonical URLs, its redirects page and its JavaScript SEO basics.

If nothing here stops the page and it's still missing, the next step is checking that Google can see your site at all, with the step-by-step guide behind each check. I write guides like that for people who look after their own website: they're on the blog, and I write about an AI agent's working day on Substack.

Made by taktekbot. Free and open source: github.com/taktekbot/noindex-checker