Skip to content
← All articles Google 403 Forbidden: Why Googlebot Is Blocked and How to Fix It

Google 403 Forbidden: Why Googlebot Is Blocked and How to Fix It

A 403 to Googlebot drops the page from the index and does nothing for crawl rate. How to find those URLs and what to return instead, from Google documentation.

A 403 returned to Googlebot is not a security setting that worked — it is a page leaving the index. Google states the position plainly in the page indexing report: HTTP 403 means the user agent supplied credentials and was refused, but "Googlebot never provides credentials, so your server is returning this error incorrectly". The page will not be indexed.

That makes 403 different from most crawl problems. It is almost never a deliberate decision about that page. It is a firewall rule, a bot filter or a rate limit that caught Googlebot by accident, and the site owner finds out weeks later when traffic has already gone.

What Google does with a 403

The crawling documentation treats the whole 4xx family the same way, with one exception: "All 4xx errors, except 429, are treated the same: Google crawlers inform the next processing system that the content doesn't exist." For Search that has two consequences, both stated in the same document:

  • "the indexing pipeline removes the URL from the index if it was previously indexed";
  • "Any content Google receives from URLs that return a 4xx status code is ignored" — so whatever your error page says, it is not read as content.

There is no grace period described, and no partial state. A page that answers 403 is a page Google treats as gone. The rest of the 4xx family behaves identically: 401, 404, 410 and 411 all end at the same place, which is why the practical fix is the same whichever of them your server returns.

403 does not slow crawling — that is the mistake

Most 403s to Googlebot exist because someone wanted less bot traffic. The documentation addresses exactly that: "Don't use 401 and 403 status codes for limiting the crawl rate. The 4xx status codes, except 429, have no effect on crawl rate."

So the block does nothing to the load and everything to the index. The documented signal for an overloaded server is 429, which Google's crawlers "treat as a signal that the server is overloaded", and it is handled as a server error rather than a missing page. If the goal is a quieter server, 429 is the answer and 403 is the damage.

Why your browser sees the page and Googlebot does not

This is what makes a 403 hard to notice: you open the URL, it works. The rules that produce it are usually conditional, and the condition is almost never "this URL".

What triggers itWhy Googlebot gets caught
WAF or bot-protection ruleThe rule matches a user agent or a request pattern, and crawler traffic looks like the pattern it was written against.
Rate limiting by user agentGooglebot requests many URLs in a short window — exactly the shape a naive rule treats as abuse.
Geo-blockingGoogle crawls primarily from US IP addresses. A country allowlist that does not include them blocks the crawler, not the audience.
Hotlink or referer protectionCrawler requests carry no referer, so a rule requiring one refuses them.
Staging authentication left in placeBasic auth or an IP allowlist survives a launch; the home page is public, deeper sections are not.
Paywall or members areaDeliberate, but then the pages should not be in the sitemap or the internal link graph either.

The second trap is that a rule can block requests claiming to be Googlebot while the real crawler is something you can verify. Google documents two ways to confirm a request genuinely came from its crawlers: a reverse DNS lookup, or matching the IP against the published Googlebot ranges. Blocking by user-agent string alone blocks the real crawler and lets the fake one through.

How to find every 403 on your site

Three checks, in order, from the widest to the most specific. None of them needs a paid tool.

  1. Crawl the site and read the status column. A crawl shows which URLs answer 403 right now, including the ones Search Console has not reported yet because it has not re-fetched them. Our site crawler lists the response code for every URL it reaches.
  2. Check a specific list of URLs. When you already suspect a section — a catalogue, a filter, a paginated archive — paste the list into the redirect and status checker and read the codes in bulk rather than one at a time.
  3. Read what Search Console already recorded. The page indexing report has a status named exactly "Blocked due to access forbidden (403)". Our indexing cleanup groups those URLs so you can see whether it is a template, a folder or the whole site.

To prove a specific URL is refused to Googlebot and not to you, run a live test in the URL Inspection tool. It fetches the page as Google and reports the response it actually got, which settles the "but it opens for me" argument in one click.

Fixing it by cause

The fix is never "return 200 to everything". It is matching the response to the intent for that URL.

SituationCorrect response
The page should be in searchAllow the crawler. Verify the request properly instead of filtering by user agent.
The server is genuinely overloaded429, not 403 — it is the documented overload signal and it is treated as a server error.
The page should never be crawledDisallow it in robots.txt, so the crawler does not spend requests on it at all.
The page is gone for good410, or 404 — both remove it, and both say something truthful.
The content moved301 to the new location.
The section is private by designKeep the 403, but remove the URLs from the sitemap and from internal links.

After the fix, request indexing for a handful of the affected URLs rather than all of them: the crawler returns on its own once the pages answer normally, and the rest of the status-code decisions are collected in our reference of HTTP status codes.

What to watch afterwards

A 403 that lasted weeks does not reverse the day you fix it. The pages have to be re-crawled, re-processed and re-indexed, and nothing in Google's documentation promises a timescale for that. What you can watch is the shape: the count under "Blocked due to access forbidden (403)" should fall, and impressions for the affected URLs should return before clicks do. If the count does not move at all after a fortnight, the rule is still firing for the crawler even though it no longer fires for you — and that is back to the verification step above. The reports for all of it live in the Search Console tools.

FAQ

Short answers to the most common questions from this article.

What does Google 403 Forbidden mean?
The server understood the request and refused it. For Googlebot specifically, Google says the response is returned incorrectly, because Googlebot never provides credentials — so a 403 cannot be the answer to a failed authorisation that never happened.
Does a 403 remove my page from Google?
Yes. The crawling documentation states that for a 4xx response the indexing pipeline removes the URL from the index if it was previously indexed, and any content returned with that status is ignored.
Can I use 403 to reduce how often Google crawls my site?
No, and Google says so directly: do not use 401 and 403 to limit the crawl rate, because 4xx codes other than 429 have no effect on it. The documented signal for an overloaded server is 429.
Why does the page open in my browser but return 403 to Googlebot?
The rule that produces it is conditional — on user agent, request rate, country or a missing referer. Your browser does not match the condition and the crawler does. A live URL inspection shows the response Google actually received.
How do I check that a request really came from Googlebot?
Google documents two methods: a reverse DNS lookup on the requesting IP, or matching the IP against the published Googlebot ranges. Filtering by the user-agent string alone blocks the real crawler and admits anything pretending to be it.
How long before the pages come back after the fix?
Google publishes no timescale. The pages have to be re-crawled and re-processed first, so impressions return before clicks. If the 403 count in the page indexing report does not fall within a couple of weeks, the rule is still firing for the crawler.