A 403 returned to Googlebot is not a security setting that worked — it is a page leaving the index. Google states the position plainly in the page indexing report: HTTP 403 means the user agent supplied credentials and was refused, but "Googlebot never provides credentials, so your server is returning this error incorrectly". The page will not be indexed.
That makes 403 different from most crawl problems. It is almost never a deliberate decision about that page. It is a firewall rule, a bot filter or a rate limit that caught Googlebot by accident, and the site owner finds out weeks later when traffic has already gone.
What Google does with a 403
The crawling documentation treats the whole 4xx family the same way, with one exception: "All 4xx errors, except 429, are treated the same: Google crawlers inform the next processing system that the content doesn't exist." For Search that has two consequences, both stated in the same document:
- "the indexing pipeline removes the URL from the index if it was previously indexed";
- "Any content Google receives from URLs that return a 4xx status code is ignored" — so whatever your error page says, it is not read as content.
There is no grace period described, and no partial state. A page that answers 403 is a page Google treats as gone. The rest of the 4xx family behaves identically: 401, 404, 410 and 411 all end at the same place, which is why the practical fix is the same whichever of them your server returns.
403 does not slow crawling — that is the mistake
Most 403s to Googlebot exist because someone wanted less bot traffic. The documentation addresses exactly that: "Don't use 401 and 403 status codes for limiting the crawl rate. The 4xx status codes, except 429, have no effect on crawl rate."
So the block does nothing to the load and everything to the index. The documented signal for an overloaded server is 429, which Google's crawlers "treat as a signal that the server is overloaded", and it is handled as a server error rather than a missing page. If the goal is a quieter server, 429 is the answer and 403 is the damage.
Why your browser sees the page and Googlebot does not
This is what makes a 403 hard to notice: you open the URL, it works. The rules that produce it are usually conditional, and the condition is almost never "this URL".
| What triggers it | Why Googlebot gets caught |
|---|---|
| WAF or bot-protection rule | The rule matches a user agent or a request pattern, and crawler traffic looks like the pattern it was written against. |
| Rate limiting by user agent | Googlebot requests many URLs in a short window — exactly the shape a naive rule treats as abuse. |
| Geo-blocking | Google crawls primarily from US IP addresses. A country allowlist that does not include them blocks the crawler, not the audience. |
| Hotlink or referer protection | Crawler requests carry no referer, so a rule requiring one refuses them. |
| Staging authentication left in place | Basic auth or an IP allowlist survives a launch; the home page is public, deeper sections are not. |
| Paywall or members area | Deliberate, but then the pages should not be in the sitemap or the internal link graph either. |
The second trap is that a rule can block requests claiming to be Googlebot while the real crawler is something you can verify. Google documents two ways to confirm a request genuinely came from its crawlers: a reverse DNS lookup, or matching the IP against the published Googlebot ranges. Blocking by user-agent string alone blocks the real crawler and lets the fake one through.
How to find every 403 on your site
Three checks, in order, from the widest to the most specific. None of them needs a paid tool.
- Crawl the site and read the status column. A crawl shows which URLs answer 403 right now, including the ones Search Console has not reported yet because it has not re-fetched them. Our site crawler lists the response code for every URL it reaches.
- Check a specific list of URLs. When you already suspect a section — a catalogue, a filter, a paginated archive — paste the list into the redirect and status checker and read the codes in bulk rather than one at a time.
- Read what Search Console already recorded. The page indexing report has a status named exactly "Blocked due to access forbidden (403)". Our indexing cleanup groups those URLs so you can see whether it is a template, a folder or the whole site.
To prove a specific URL is refused to Googlebot and not to you, run a live test in the URL Inspection tool. It fetches the page as Google and reports the response it actually got, which settles the "but it opens for me" argument in one click.
Fixing it by cause
The fix is never "return 200 to everything". It is matching the response to the intent for that URL.
| Situation | Correct response |
|---|---|
| The page should be in search | Allow the crawler. Verify the request properly instead of filtering by user agent. |
| The server is genuinely overloaded | 429, not 403 — it is the documented overload signal and it is treated as a server error. |
| The page should never be crawled | Disallow it in robots.txt, so the crawler does not spend requests on it at all. |
| The page is gone for good | 410, or 404 — both remove it, and both say something truthful. |
| The content moved | 301 to the new location. |
| The section is private by design | Keep the 403, but remove the URLs from the sitemap and from internal links. |
After the fix, request indexing for a handful of the affected URLs rather than all of them: the crawler returns on its own once the pages answer normally, and the rest of the status-code decisions are collected in our reference of HTTP status codes.
What to watch afterwards
A 403 that lasted weeks does not reverse the day you fix it. The pages have to be re-crawled, re-processed and re-indexed, and nothing in Google's documentation promises a timescale for that. What you can watch is the shape: the count under "Blocked due to access forbidden (403)" should fall, and impressions for the affected URLs should return before clicks do. If the count does not move at all after a fortnight, the rule is still firing for the crawler even though it no longer fires for you — and that is back to the verification step above. The reports for all of it live in the Search Console tools.