How to find old pages of your website that still show up in Google, and remove them properly
A page you forgot is still a page a search engine reads. Here is how to find yours and retire them without breaking anything.
To find old pages that still show up in Google, search site:yourdomain.com with words like "price", "hours" or a past year, and export the indexed pages from Google Search Console. For each old page, update it, redirect it to the page that replaced it, or delete it so it returns a 404. Then use Google's Removals tool to hide it while Google catches up.
Old pages matter more than they used to. Search engines show them, and AI assistants that search the web read them and repeat what they say. If ChatGPT quotes your 2023 prices, there's usually a 2023 page behind it. This post is the cleanup that follows checking what assistants say about you.
Step 1: find the old pages
Use more than one way. Each one finds pages the others miss.
Search Google with operators
Type these into Google, with your own domain:
site:yourdomain.com price
site:yourdomain.com hours
site:yourdomain.com 2023
site:yourdomain.com offer OR promotion OR holiday
site:yourdomain.com before:2024
before: finds pages last updated before a date, and takes a year or a full date like before:2024-06-01. Google documents it on its Refine searches page, along with site:, - and after:. The date is Google's idea of when the page changed, so treat it as a hint, not a fact.
To look at one stretch of time, combine the two: site:yourdomain.com after:2021-01-01 before:2023-01-01 shows pages Google thinks were last updated in 2021 or 2022. That's handy when you know roughly when the old prices or the old address were on the site.
Search for the things that go out of date: prices, hours, addresses, branch names, staff pages, menus, events, job ads, and the names of products you no longer sell.
Export the list from Search Console
- In Google Search Console, open Indexing → Pages.
- Click View data about indexed pages.
- Click Export and open the file in a spreadsheet.
The table stops at 1,000 rows (source). Most small sites fit. Read the addresses top to bottom. Old pages give themselves away by their paths: /2022-menu, /summer-sale, /branch-old-street.
Do the same in Bing Webmaster Tools. Bing feeds Microsoft Copilot, and its index is not a copy of Google's.
Check what still gets visits
Before you delete anything, open Search Console's Performance report, click Pages, and look up each old page. A page that still gets clicks has people arriving on it. It needs a redirect, not a dead end.
Step 2: decide what each page becomes
Every old page gets one of four outcomes. Write it in a new column next to the address.
- Update when the page is still useful but some facts changed. Fix the facts, keep the address. This keeps any links pointing to it.
- Redirect when a newer page does the same job: the 2023 price list becomes the current price list, the closed branch points to the nearest open one. Use a permanent redirect (301 or 308). Google treats it as a strong signal that the new page should take the old one's place (source). With more than a handful, list them in two columns and paste them into the redirect map builder: it writes the rules for Apache, nginx, Netlify, Cloudflare or Vercel and catches chains and loops.
- Delete when nothing replaced it: an expired offer, an event that happened, a product you'll never sell again. The server must answer with a 404 or 410 status.
- Noindex when some people still need the page but searchers shouldn't land on it, like an old terms-and-conditions page that existing customers signed. Add
<meta name="robots" content="noindex">to its<head>. Old price lists are often PDFs, which have no<head>. For those, the server sends the headerX-Robots-Tag: noindexinstead (Google and Bing both read it). Or just delete the PDF.
404 or 410 doesn't matter much. Google says it treats all 4xx codes except 429 the same way: the page is dropped from the index if it was there (source). Pick whichever your website builder makes easy.
Step 3: check that it actually worked
Website builders get this wrong more often than you'd think. A "deleted" page can still answer with a 200 status and a friendly "page not found" message. Google calls that a soft 404, and it stays a problem (source).
Check each changed address. In a terminal, on Mac or Linux:
curl -sI https://yourdomain.com/summer-sale | head -1
# HTTP/2 404 deleted, good
# HTTP/2 301 redirected, good
# HTTP/2 200 still there: not done
For a redirect, also check where it goes:
curl -sIL https://yourdomain.com/2023-prices | grep -i '^location'
On Windows, PowerShell has curl.exe built in but no head or grep. Use these instead:
curl.exe -sI https://yourdomain.com/summer-sale | Select-Object -First 1
curl.exe -sIL https://yourdomain.com/2023-prices | Select-String '^location'
No terminal? Use Search Console's URL Inspection: paste the address in the bar at the top, then click Test live URL. It shows what Google gets back right now.
Step 4: hide it now, while Google catches up
Google drops a deleted page only when it crawls that address again. That can take weeks for a page nobody links to. If the old page is doing harm, like quoting a wrong price, hide it in the meantime:
- Google: Search Console → Indexing → Removals → New request. This hides the page from search for about six months. It does not delete it. Google's own words: the tool "provides only a temporary removal" (source). The 404, redirect or noindex from step 2 is what makes it permanent.
- Bing: Bing Webmaster Tools' Block URLs. Also temporary; a block lasts 90 days unless you extend it. Bing's page on removing a page from Bing or Copilot for good lists the same permanent fixes: a 404 or 410, or noindex.
If you already use IndexNow, submit the changed and deleted addresses there too. It tells Bing and other engines that use it to look again.
When the old page isn't yours
Sometimes the old fact sits on someone else's site: a directory, a news article, a blog. If that page is gone or has already been fixed but the old version still shows in search, anyone can ask for a refresh:
- Google: the Refresh Outdated Content tool.
- Bing: the Content Removal tool.
If the page is live and still wrong, these tools can't help. Ask the site to update it. Most directories let owners claim and edit their listing.
Mistakes to skip
- Blocking old pages in robots.txt. It feels like removal. It isn't. If Google can't fetch the page, it never sees the noindex, and "the page can still appear in search results" (source). Let the crawler in so it can read the 404 or the noindex. To see whether a line in your robots.txt still blocks a page, paste the file and the address into the robots.txt tester.
- Redirecting everything to the home page. It's one rule instead of twenty, and it helps nobody. Someone who clicked "2023 price list" wants prices, not your welcome banner. Redirect each page to its closest match, and delete the ones that have none.
- Using a temporary redirect. 302 and 307 tell Google the move might be undone, so it may keep showing the old address. Use 301 or 308 for pages that are gone for good.
- Panicking at the 404 count. After a cleanup, Search Console's Not found (404) list grows. For pages you deleted on purpose, that is the report doing its job, not an error to fix.
- Forgetting the links. Search your own site for links to the pages you removed and change them. Your menus and old blog posts are often the reason a dead page keeps getting crawled.
The checklist
- Search
site:yourdomain.comwith "price", "hours", past years, andbefore:. - Export indexed pages from Search Console (and Bing Webmaster Tools) into a spreadsheet.
- Look up each old page in the Performance report before you touch it.
- Mark each one: update, redirect (301), delete (404/410) or noindex.
- Make the change, then check the status code with
curl -sIor URL Inspection. - Hide harmful pages with Google's Removals tool and Bing's Block URLs while you wait.
- Fix your own links to removed pages.
- Search
site:again in a few weeks, and repeat once a year.
taktekbot