robots.txt: Crawling, Not Indexing
robots.txt controls crawling, not indexing. The difference is not academic: a single line in this file can zero out a site's visibility, and using it to remove a page from search usually fails — and the documentation explains why.
The lesson follows Google's specification: the introduction to robots.txt and how Google interprets the file. The second is rarely read, and it describes the behaviour everyone eventually runs into: caching, server responses and limits.
/robots.txt. The cause is in the first rules.Disallow: /
# temporarily block bots
# remember to revert before launch
What robots.txt does
robots.txt is a text file at the root of a host: https://example.com/robots.txt. It tells well-behaved crawlers which URLs they may and may not crawl. Its main purpose, as the documentation puts it, is to avoid overloading your site with requests.
noindex or password protection.The difference between crawling and indexing is what people trip over most. Disallow stops a crawler from visiting a page. But if other sites link to it, Google may index the URL without visiting it: the result appears without a description. That is documented plainly, and it means something simple — blocking a page from crawling did not remove it from search, it only removed Google's ability to know what is on it.
The same fact creates the main trap: if a page is blocked in robots.txt, the crawler never sees the noindex on it. Two prohibitions together cancel each other out. The glossary covers the terms: robots.txt and noindex.
Basic Syntax
User-agent: *
Disallow: /admin/
Disallow: /cart/
Allow: /
| Line | Meaning | Comment |
|---|---|---|
User-agent: * | Rules for all crawlers | You can target a specific crawler such as Googlebot or Googlebot-Image. |
Disallow: /admin/ | Do not crawl URLs starting with /admin/ | Paths are case-sensitive: /Private/ and /private/ are different. |
Allow: / | Allow crawling of the rest of the site | Often optional, but it makes intent obvious for beginners. |
Sitemap: https://example.com/sitemap.xml | Advertise the XML sitemap location | This line can sit outside rule groups. |
Scope and limits
Three things from the specification that save hours of debugging.
| Rule | What it means in practice |
|---|---|
| The file applies to its own protocol, host and port | The robots.txt for https://example.com does not govern http://example.com or the subdomain shop.example.com — each has its own file |
| Google honours at most 500 KiB | Anything past that limit is ignored. A huge file with thousands of lines is a sign the rules need consolidating |
| Paths are case-sensitive | /Private/ and /private/ are different rules |
The full syntax, with directive examples and the order rules are applied in, is in the guide to creating the file.
Dangerous Lines
User-agent: *
Disallow: /
This combination blocks crawling for the whole site. It is sometimes used on development environments and accidentally shipped to production.
User-agent: *
Disallow:
An empty Disallow means nothing is blocked. It is not an error, but it can look ambiguous. On a public site, a sitemap-only file or a clear comment is often easier to read.
User-agent: *
Disallow: /*?sort=
Disallow: /*?filter=
Patterns can limit noisy parameters, but test them against real URLs. One broad mask can accidentally block useful filters, pagination, or listings that should receive search traffic.
What happens when the file is unavailable
Nobody thinks about this until it happens. And it happens regularly: the server went down, a deploy broke the route, the CDN returned an error. Google's behaviour is described in the specification and differs by response.
| Server response | What Google does |
|---|---|
| 4xx, except 429 | Treats it as if the file did not exist and assumes no crawl restrictions. A 404 on robots.txt is a normal situation, not a problem |
| 5xx | For the first 12 hours it stops crawling the site while still trying to fetch the file. After that, for up to 30 days, it uses the last good version |
| Redirect | Follows at least five hops, then treats it as a 404 |
/robots.txt for an hour halts crawling of the entire site, not just that file.And one more thing that makes a fix look like it failed: Google generally caches the file's contents for up to 24 hours. Removing a stray Disallow: / will not take effect instantly — that is normal, not a reason to edit the file again.
The noindex rule
If the goal is to remove a page from search, do not start with Disallow. Leave the page crawlable and add:
<meta name="robots" content="noindex">
For files with no HTML — PDFs, images — the X-Robots-Tag header does the same job. Both methods and their syntax are in the documentation on the robots meta tag, and the overall removal procedure in the guide to blocking indexing.
The order matters: first let Google see the noindex and drop the URL from the index, then restrict crawling if you still need to. Done the other way round, the page stays in the results — without a description and with no way to fix it.
How to check the file
The check takes five minutes and comes in three forms, from fast to precise.
- Open
/robots.txtin a browser. Does it return 200? Is there a strayDisallow: /on the production site? - Look at the robots.txt report in Search Console: which version Google currently sees and when it fetched it. Parsing errors show up here too.
- Check a specific URL with URL Inspection. If the reason is "Blocked by robots.txt", you have found the problem. The URL auditor does this in bulk.
How many pages dropped out of the index and why is in the page indexing report; the gap between what exists on the site and what Search sees is the coverage gap. The glossary covers the terms: crawl budget and index coverage.
A Normal robots.txt for a Small Site
User-agent: *
Disallow: /admin/
Disallow: /cart/
Disallow: /checkout/
Disallow: /search
Sitemap: https://example.com/sitemap.xml
the file is available at
/robots.txt and returns 200;production does not contain accidental
Disallow: /;important templates are open to Googlebot;
sitemap is listed as a full URL;
GSC URL Inspection does not show "Blocked by robots.txt" for important pages.
Practice: check your own robots.txt
Five steps, ten minutes:
- Open
https://your-domain/robots.txt. Read the whole file — it is usually twenty lines, half of which stopped being needed long ago. - Write out every
Disallowline and say what exactly it blocks. If you cannot, the line is a candidate for deletion. - Check that CSS, JS and images are not blocked. Without them Google understands the page less well — look at it as a crawler does with the page analyzer.
- Make sure no page carrying
noindexis also blocked in robots.txt. That is the most common pair of mistakes, and it leaves the page in the results permanently. - Check the subdomains: each has its own file. A crawl covers the whole site, and sitemap monitoring the state of the sitemap.
How to tell the file is in order
The sign is not that the file is "correct" but that you can explain every line in it. robots.txt is a rare case where a short file beats a thorough one: the fewer rules, the fewer ways to block something you needed.
Three questions to check yourself. Do you know what each line in your file blocks? Do you understand the difference between "do not crawl" and "do not index"? Would you notice a Disallow: / reaching production the same day?
What not to expect: that robots.txt removes a page from search or protects a private area. The file is public — anyone can open it — and every path listed in it is visible to everyone. Secret directories are not hidden by it, they are advertised.
What comes next
This is the last lesson of the SEO Foundations course. If you did not start at the beginning, Google Search Console is worth going back to: nearly every check in this lesson happens there. The previous lesson, on platforms and tools, is here.
Disallow: /, kept only utility sections blocked, and added the sitemap. Google can crawl the important pages again.