Does my small business website need a robots.txt? Probably not
Most advice says every site needs a robots.txt. For a site with a few dozen pages, Google's own help says having none is fine. Here is when that changes.
If your website has a few dozen pages and you want all of them in Google, you don't need a robots.txt file. You can leave it empty, or not have one at all.
That isn't my opinion. Google's Search Console help says it plainly: "Not having a robots.txt file is fine, and means that Google can crawl all URLs on your site" (source).
What matters more is how your server answers when a crawler asks for the file. On 8 October 2026 I asked 35 websites for their robots.txt. 33 gave a normal answer: 31 sent a file, 2 said "not found", and both of those are fine. The other 2 sent a bot check page instead of a robots.txt. From the outside, that's the one result you can't read, and it's covered below.
I'm saying all this even though I built a robots.txt tester. Most people who reach for the file don't need it. Rules copied from a forum often do more harm than having no file at all.
Below: what the file is for, the few cases where a small site needs a rule, the mistake almost everyone makes with it, and the one thing about robots.txt that can stop Google from crawling your site.
What robots.txt actually does
It's a plain text file at yourdomain.com/robots.txt. It tells crawlers which addresses on your site they may request. That's all it does.
Google describes it this way: it "is used mainly to avoid overloading your site with requests; it is not a mechanism for keeping a web page out of Google" (source).
So the file answers one question: should a crawler fetch this address? It doesn't answer "should this page appear in search results?" Those are different questions, and mixing them up causes most of the problems below.
When there's no file, Google treats the site as having no restrictions. The same goes for an empty file, or one that only says User-agent: * followed by an empty Disallow:. All three mean "crawl everything".
If you're on Wix, Squarespace or a similar host
Your host already serves a robots.txt for you. Google's intro page notes that on a CMS like Wix or Blogger, you "might not need to (or be able to)" edit the file directly. Use the host's settings to hide a page from search engines instead.
To find that setting, search for your host's name plus "hide page from search engines". Google suggests the same search.
If your site is behind Cloudflare, your robots.txt may have lines you never wrote. With its managed robots.txt on, Cloudflare adds a block at the top, between # BEGIN Cloudflare Managed content and # END Cloudflare Managed Content, asking AI training bots such as GPTBot and ClaudeBot to stay away. Search bots aren't on its list. On the Free plan, if your site has no robots.txt of its own and the managed setting is off, Cloudflare serves a page of comments about "content signals" instead. It has no rules, so it blocks nothing. Both are switched on and off in Cloudflare, not on your site (Cloudflare's page).
The few cases where a small site needs a rule
Your site makes endless addresses. An online shop with filters can turn 40 products into thousands of URLs: ?colour=red&size=m&sort=price, in every combination. The same goes for a site search that makes a new page for every query. Google calls these "unimportant or similar pages", and blocking them is a fair use of the file. A rule like this stops crawlers from requesting any address with a query string:
User-agent: *
Disallow: /*?
Be careful with it. If your real product pages also use ? addresses, like /product?id=12, this rule blocks them too. Test it against your own URLs first.
Your server struggles when crawlers visit. If your host's logs show the site slowing down while bots crawl it, robots.txt is the tool Google intends for that. On a site of a few dozen pages, this almost never happens.
You want to keep images or videos out of Google. For media files, a Disallow rule does keep them out of search results. Google's intro page lists this as a separate case from web pages.
If none of these describes your site, you have nothing to add.
The mistake: using Disallow to hide a page
This is the one to remember. A page blocked in robots.txt can still show up in Google.
Google's words: "A page that's disallowed in robots.txt can still be indexed if linked to from other sites." Google won't read the page, but it can still list the address, with your link text, and no description under it.
It gets worse if you combine the two methods. Say you add a noindex tag to a page and also block it in robots.txt. Google can't fetch the page, so it never sees the tag, and the page can stay in the results. Google explains this on its noindex page.
So, by goal:
- A page you don't want in search results (a thank-you page, an old offer): add
<meta name="robots" content="noindex">and leave it crawlable. - A copy of your site you're still building: put a password on it. Most hosts have a setting for this. A robots.txt rule only asks well-behaved crawlers to stay away, and the address can still get listed.
- Something private (invoices, customer files): a password, always. Google says robots.txt "cannot enforce crawler behavior", and anyone can open your robots.txt and read the list of addresses you'd rather they didn't visit.
WordPress changed its own setting for exactly this reason. Since version 5.3, "Discourage search engines from indexing this site" adds a noindex tag instead of a robots.txt block (WP Tavern). If your WordPress site isn't showing up, go to Settings > Reading first and make sure that box is unticked.
The thing that does matter: what your server answers
A missing robots.txt is harmless. A broken one isn't. What decides it is the status code your server sends back when Google asks for /robots.txt. Google's robots.txt rules say:
- 200: Google reads the file and follows it.
- 404, or any other 4xx except 429: Google acts as if there's no file and crawls everything. This is the safe answer for "I don't have one".
- 5xx, a timeout, or a DNS failure: for the first 12 hours, Google stops crawling the whole site and keeps retrying the file. After that, for up to 30 days, it uses the last version it read. If it never had one, it assumes no restrictions.
The third case is the trap. A server that errors on /robots.txt, or a firewall that times out on Google's request, pauses crawling of every page on the site. And you won't see it, because the file looks fine in your browser.
Check what your server answers. On a Mac or Linux terminal:
curl -sL -o /dev/null -w "%{http_code}\n" https://yourdomain.com/robots.txt
On Windows, in PowerShell:
curl.exe -sL -o NUL -w "%{http_code}" https://yourdomain.com/robots.txt
You want 200 (you have a file) or 404 (you don't). Both are fine. A 500, 502, 503, or 000 (curl got no answer at all) is the one to fix, usually by asking your host.
The 35 sites I checked on 8 October were 24 online shops and 11 well-known sites, and this is the command I ran. 31 answered 200 and 2 answered 404. The other two answered with a bot check, not a robots.txt file. One sent 403 and a Cloudflare page titled "Just a moment...". The other sent 429 and a page titled "Vercel Security Checkpoint". Their headers said the same thing (cf-mitigated: challenge and x-vercel-mitigated: challenge).
So if you get a 403 or a 429, look at what came back before you act on the list above:
curl -sL https://yourdomain.com/robots.txt | head -5
On Windows: curl.exe -sL https://yourdomain.com/robots.txt | Select-Object -First 5. If you see HTML with a title like "Just a moment..." or "Security Checkpoint", your firewall stopped curl. That doesn't tell you what Google got. Open the Search Console report below: it shows what Google itself fetched. "Fetched" means only curl was stopped.
The command follows redirects and prints the status of the last answer. That matters if you type http:// or leave out the www: the first answer is then a 301, which tells you nothing about the file. Google follows redirects too: at least five hops, then it stops and treats the file as missing (source).
No terminal? In Search Console, go to Settings, then robots.txt, or open the report directly. "Fetched" means Google read your file. "Not Fetched – Not found (404)" means you don't have one, which is fine. "Not Fetched – Any other reason" is the one to look into. The report only works for a Domain property or a URL-prefix property without a path.
If you want one line, make it the sitemap
The one line worth adding to a small site's robots.txt doesn't block anything. It tells crawlers where your sitemap is:
Sitemap: https://yourdomain.com/sitemap.xml
Google accepts this as one way to submit a sitemap, and it helps crawlers whose webmaster tools you never signed up for. If you've already submitted the sitemap in Search Console and Bing Webmaster Tools, it's optional.
Whether to block AI crawlers is a separate choice, and a real one. If you want your business to show up in ChatGPT search, don't block OAI-SearchBot. I go through that in what to put on your site so AI assistants can answer questions about it.
If you already have a robots.txt
Open it. If it's more than a few lines and you didn't write them, check that none of them blocks a page you care about. Look hardest at Disallow: / under User-agent: *, which blocks the whole site. It's often left over from when the site was being built.
The quickest way to check is to paste the file and a few of your page addresses into my robots.txt tester. It shows, for each address, whether Googlebot may crawl it and which line decided. It runs in your browser and sends nothing anywhere.
If the rules turn out to be ones you don't need, delete them. An empty file is a fine file.
After you change it, don't expect to see the new version at once if your site is behind Cloudflare. When I checked a Cloudflare-fronted site on 8 October 2026, its robots.txt came back with cache-control: max-age=14400: browsers and crawlers may keep a copy for up to four hours. Test again after a wait, not the moment you save.
The short version
- A small site that wants every page in Google needs no robots.txt rules. No file, or an empty one, is fine.
- To keep a page out of Google, use
noindex, not Disallow. For anything private, use a password. - Never block a page in robots.txt and also noindex it: Google can't see the tag.
- Check that
/robots.txtanswers 200 or 404, never 5xx or a timeout. - The one useful line for a small site is
Sitemap:.
Sources: Google's Introduction to robots.txt (last updated 2025-12-10), How Google interprets the robots.txt specification (status codes, 12 hours, 30 days; last updated 2026-08-31), Search Console's robots.txt report help, and Google's noindex page.
taktekbot