Indexing issues: a technical audit from sitemap to URL Inspection
A technical indexing audit checks whether the pages that should be in Search can be there, and why not when they are missing. This playbook turns one coverage total into a prioritized queue, maps statuses to the next check and lists mistakes that look like fixes.
✓ Verified against Google documentation · 4 October 2026
What a technical indexing audit is
A technical indexing audit answers one question: are the pages that should be in Search eligible to be there, and if not, why? It starts from the Page indexing report, which groups the URLs Google knows about by state (Page indexing report). A status is a queue, not a root cause: the same label can come from different mechanisms, and many labels describe exclusions you chose.
Google is explicit that not being indexed is often fine: robots.txt rules, noindex tags, duplicates and removed pages are normal reasons. So an audit does not chase “everything indexed”. It chases “every important page indexed and every exclusion explained”. If the lesson on how indexing works is new to you, read it first.
Four gates for every URL
Do not prescribe a site-wide fix until a URL has passed through these questions.
- Search intent. Should this canonical URL appear in Search, or is it a deliberate redirect, duplicate, filter, private page or removal?
- Fetch control. Check the HTTP response, robots.txt reachability, noindex and whether Google can fetch the resources needed to understand the page.
- Canonical consistency. Compare redirects, the declared canonical, the Google-selected canonical, internal links and sitemap inclusion.
- Page value. Confirm the rendered page is distinct, useful and reachable through internal links; a status label alone does not prove a quality diagnosis.
Turning one coverage total into a queue
A neutral release sitemap with 240 URLs. Review finds 38 intentional exclusions; of the 202 important URLs, 185 are indexed and 17 need investigation.
| Group | URLs | Decision |
|---|---|---|
| Submitted inventory | 240 | starting set |
| Expected exclusions | 38 | accept and document |
| Important working set | 202 | 240 − 38 |
| Indexed important URLs | 185 | monitor |
| Unexpected exclusions | 17 | inspect and assign an owner |
Working-set coverage = 185 ÷ 202 × 100 = 91.6%
Investigation queue = 202 − 185 = 17 URLs
The percentage is an internal control, not a Google target. The deliverable is the 17-URL queue split by template and business value, with representative inspection evidence.
The eight-step triage
Work from a defined URL set and keep the evidence for each decision.
- Define the working set. Use current canonical sitemap URLs, a release batch or a business-critical directory, not the whole discovered inventory. Google asks for absolute, canonical URLs that you want in results (build a sitemap).
- Open the Page indexing reasons. Record indexed and not-indexed counts for that set and export examples where possible.
- Remove expected exclusions. Document intentional redirects, duplicates, noindex pages and removals so they do not enter the fix queue.
- Sample by template and value. Inspect important examples from every unexpected status, directory and template (URL Inspection tool).
- Compare indexed and live state. Record the coverage reason, last crawl, fetch state, robots state and declared versus Google-selected canonical.
- Confirm the root mechanism outside the label: server response, rendered HTML, directives, internal links and sitemap.
- Ship one cause-specific fix. Assign an owner and change only what the evidence supports; validate one representative URL before applying it to a template.
- Request recrawl and monitor. Use Request indexing for a few URLs or a sitemap for many; neither guarantees inclusion or timing (ask Google to recrawl).
Map the evidence to the next check
| Status | Next check |
|---|---|
| Discovered – currently not indexed | Google found the URL but has not crawled it; typically it expected the crawl to overload the site. Confirm the URL belongs in Search, then check discovery paths, internal links, server speed in the Crawl Stats report and low-value URL expansion (Crawl Stats report) |
| Crawled – currently not indexed | Google crawled the page but did not index it, and says you need not resubmit it. Compare live rendering, canonical signals, duplication and usefulness; Google does not name one universal cause |
| Soft 404 | if the content is gone return a real 404; if the page should exist, make sure it returns meaningful content |
| Excluded by noindex | keep intentional exclusions; for an important URL remove the directive and leave crawling allowed so Google can see the change |
| Blocked by robots.txt | fix an accidental disallow where crawling is needed; robots.txt is not a mechanism for keeping a page out of Google (robots.txt introduction) |
| Duplicate or different canonical | inspect both canonicals and align redirects, rel=canonical, internal links and sitemap signals (consolidate duplicate URLs) |
Mistakes that look like fixes
- Blocking a page in robots.txt and adding noindex. For noindex to work the page must not be blocked: if the crawler cannot access it, it never sees the rule and the page can still appear in results (block indexing with noindex).
- Using robots.txt to hide a page. It manages crawl traffic; use noindex or password protection to keep a page out of Google.
- Requesting indexing again and again. There is a quota and repeated requests for the same URL do not make crawling faster.
- Listing non-canonical or relative URLs in the sitemap. List absolute, canonical URLs you want in results.
- Treating the live test as the indexed state. It checks the current retrievable version, which can differ from the indexed one and does not guarantee indexing.
What this triage cannot prove
- No 100% target. Google does not expect every known URL to be indexed; healthy exclusions are normal.
- No single-cause labels. A report reason summarizes Google’s state and does not replace server, rendered-page and canonical checks.
- No fixed recovery deadline. Crawling can take days to weeks and depends on many systems.
- A small site may not need it. Google says that with fewer than about 500 pages you probably do not need the Page indexing report;
site:checks are enough.
The indexing queue
For each unexpected URL record: purpose, status, HTTP response, robots and noindex state, declared and Google-selected canonical, last crawl, page value, owner and next action. Validate one representative fix before applying it to a template. If a result is unclear, record “insufficient evidence” and change nothing.
Indexing issues, answered
Why are my pages not indexed?
Open URL Inspection and read the reason: blocked, noindex, duplicate, 404, soft 404, server error or simply not crawled yet. Then ask whether the page should be in Search at all.
Should every page be indexed?
No. Exclusions by robots.txt, noindex, duplication or removal are normal.
Can I use robots.txt and noindex together?
No: a page blocked by robots.txt never shows the noindex rule to the crawler.
Does the live test mean the page is indexed?
No. It verifies the current retrievable version, not the indexed one.
How many times can I request indexing?
There is a quota, and repeated requests for one URL do not speed up crawling.
What to read next
Go back to CTR recovery or week 1 quick wins, or open the URL Inspection tool to start your own working set.