Log File Analysis
Studying server access logs to see which URLs crawlers requested, when and with what response.
Reviewed by Alexander Yarovenko · Updated: 2026-10-04
Log file analysis is the study of your web server’s access logs to see which URLs crawlers actually requested, when, and with what response. Unlike a crawl by an SEO tool, a log shows what search engine robots really did on your site. It does not show whether a page was indexed or ranked: a request is only a request.
What it is and where the boundary lies
Every request to your server is written to a log: time, address, URL, status code, size and user agent. Filtered to verified Google crawlers, that file is the only first-hand record of crawling that you own. Search Console’s Crawl Stats report is the same idea from Google’s side, but it is limited to the selected domain and is aimed at advanced users. Neither tells you what happened after the fetch.
How it differs from similar terms
| Concept | Source | What it answers |
|---|---|---|
| Log file analysis | your server | which URLs were requested and what the server answered |
| Crawl Stats report | Search Console | Google’s requests, responses and availability problems |
| Crawlability audit | a tool that crawls your site | what a robot could reach, not what Google did |
| Crawl budget | a concept from Google’s guide | how much Google can and wants to crawl |
Why it matters
For a large or frequently changing site, logs show where crawling goes: valuable pages, parameter duplicates, redirects or errors. Google says slow responses and server errors lower the crawl limit and that soft 404 pages keep being crawled and waste budget. For a small site the Crawl Stats report is described as unnecessary below roughly a thousand pages, so logs there are a diagnostic of last resort.
What to read in a log line
A typical line looks like this:
66.249.66.1 - - [03/Oct/2026:09:14:02 +0000] "GET /catalog/pans?page=2 HTTP/1.1" 200 18342 "-" "Mozilla/5.0 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)"
It records the address that asked, the time, the URL, the status (200), the size and the user agent. The user agent only claims an identity. Google’s documentation says to verify with a reverse DNS lookup of the IP, which must resolve to googlebot.com, google.com or googleusercontent.com, followed by a forward lookup that returns the same IP, or to match the IP against Google’s published lists.
How to apply it
- Get raw access logs for a period that covers several weeks, including your CDN or load balancer if one sits in front.
- Keep only requests whose IP passes the reverse-and-forward DNS check or appears in Google’s published ranges.
- Group URLs by template: category, product, article, parameter variants, files.
- Count requests and status codes per group; look for errors, long redirect chains and soft 404s.
- Compare crawled URLs with your sitemap: important URLs never requested, and requested URLs you never listed.
- Fix causes (links, redirects, response time, robots.txt rules) and repeat on the next period.
Practical example
A shop’s logs for four weeks, after verification, are grouped like this:
| Template | Finding | Action |
|---|---|---|
| /catalog/…?sort=, ?page= | most requests go to sorted variants | review internal links and canonical or robots rules |
| /product/… | new products first requested after weeks | check sitemap lastmod and internal links |
| /old-section/… | requests end in a chain of redirects | point links straight to the final URL |
Each row is a hypothesis to test, not a verdict: the log shows requests, not their effect on indexing.
Common mistakes
- Trusting the user agent: verify the IP; the string alone is not proof.
- Reading a request as indexing: a crawled page can still be left out of the index; check the page in Search Console.
- Missing the CDN layer: logs from only the origin server can hide requests answered by the cache.
- Analysing one day: crawling varies; use several weeks.
- Applying the method to a tiny site: Google says the Crawl Stats report is not needed below a thousand pages; spend the time on content.
- Counting redirect hops as one: Google’s report counts each request in a chain separately, and so do logs.
How to validate the result
After a fix, compare the same template groups for the next period: fewer requests to wasted variants, fewer error and redirect responses, earlier first requests for new pages. Confirm in the Crawl Stats report that server errors and availability problems fall. Do not expect a ranking change from crawl changes alone.
More questions
Can I get this from Search Console instead?
Partly. The Crawl Stats report shows Google’s requests, responses and availability, but only for the selected domain and in aggregate. Logs show every URL and every crawler.
Does more crawling mean better rankings?
The documentation does not say so. Crawling is a precondition for indexing, not a ranking factor.
Do I need special software?
For a few weeks of logs, a script or spreadsheet is enough. Specialised tools help with very large files.
Next practical step
Run the lesson Crawl audit, then pull one week of logs and group them by template.
Related concepts
- Crawl budget — how much Google crawls on a site.
- Crawl depth — clicks needed to reach a page.
- Crawl errors — failed requests.
- Redirect chain — several hops before the final URL.
- Google Search Console — Crawl Stats and indexing reports.