How Search Engines Work
Search engines work in four stages: crawling, rendering, indexing and ranking. Until a page has cleared the first three it cannot appear for any query at all, however well the copy is written. This lesson walks through each stage, where pages usually get stuck, and how to check your own.
Google describes the overall flow in «How Google Search works» and lists its robots in the crawler overview.
Simple answer: the search engine doesn't know your site exists yet. When you search Google, the answer is already prepared — Google has crawled the web in advance, read billions of pages, organized them, and now just picks the relevant ones. Let's break down what's happening behind the scenes.
A search engine isn't an AI chat. It's a giant pre-built database
When you type a query, Google doesn't run across the web in real time — that would be catastrophically slow. Instead, it gathers information about every page in advance into a massive database called the index. Your query then hits the index in milliseconds.
The pipeline has four sequential stages. At any one of them your site can get stuck — and then no amount of SEO will help.
Stage 1. Crawling — discovering pages
Googlebot is a "spider" program. It finds new URLs in two ways:
Which URLs the crawler visits and which it skips comes down to your sitemap and robots.txt: how Google reads a sitemap and what robots.txt does.
- By following links — it hops from a known page to anything it links to. If any indexed site links to yours, the bot eventually arrives.
- Via sitemap.xml — if you've submitted a sitemap in Google Search Console, the bot gets your URL list directly.
- Nothing links to it and it's not in the sitemap — the bot has no idea it exists
- It's buried too deep (5+ clicks from the homepage) — the bot decides it's not worth the cost
- The bot crawled it before and decided it's not worth a recrawl (thin content, duplicate)
- Your crawl budget is exhausted — the cap on how many pages Google is willing to fetch from your domain per day
Crawl budget is a real constraint. The web is too big for Google to crawl exhaustively, so it prioritizes: large authoritative sites get crawled often, small new ones much less frequently.
Stage 2. Rendering — Google actually executes the page
This step didn't exist a decade ago. Back then Google just read the raw HTML. In 2026, most sites are JavaScript-driven — React, Vue, Next.js — and the actual content only shows up after scripts run.
What Google does with JavaScript, and why content drawn by a script reaches the index later, is covered in the JavaScript SEO basics.
So Google now runs the page like a real browser (Chromium): loads HTML, executes JavaScript, waits for things to settle, and only then "snapshots" the final DOM. This step is called rendering.
Stage 3. Indexing — putting things on shelves
Google semantically analyzes the rendered page: it parses headings, text, images, structure, meta tags, and schema.org markup, extracts key concepts, and stores them in the index.
This is where Google picks a canonical URL among similar pages; the rules are in the canonicalization documentation.
The modern index is not just a text database. It's a connected web of stores:
| Store | What it holds |
|---|---|
| Main index | Metadata for every page: text, headings, links, sitemap data |
| Vertical indexes | Separate stores for images, videos, news, products — each with its own attributes |
| Knowledge Graph | A graph of entities: people, companies, places, concepts, and the relationships between them. Sources include Wikipedia, Wikidata, and schema markup |
The Knowledge Graph is why Google understands that "F1" = Formula One = motor racing, that "Elon Musk" is connected to Tesla and SpaceX, that "buy" and "sell" are close concepts. Modern search isn't word matching — it's entity matching.
⚠️ Being in the index ≠ being in the top results. The index holds billions of pages. Page 1 of the SERP holds ten. Between them sits the ranking step.
Stage 4. Ranking — picking the best answers
When a user submits a query, Google in milliseconds:
Google lists the systems involved in ranking in its guide to ranking systems.
- Identifies the query's intent (informational / transactional / local)
- Pulls relevant pages from the index
- Ranks them by hundreds of signals and assembles the SERP
In 2026, a major share of ranking is driven by machine learning. Google's key ML systems:
| System | What it does |
|---|---|
| RankBrain (2015) | Understands the meaning of rare or long-tail queries via embeddings |
| BERT (2019) | Reads each word in context of the full query (word order matters: "ticket to Boston from NYC" ≠ "ticket from Boston to NYC") |
| MUM (2021) | Connects topics across languages and content types (text + video + images together) |
| SpamBrain | Auto-detects and demotes spam, thin content, and low-quality AI-generated junk |
The main ranking signal groups:
- Relevance — does the page match the query. The largest factor by far.
- Authority — who links to you, how often you're cited in your niche (the heir to PageRank).
- EEAT — Experience, Expertise, Authoritativeness, Trustworthiness. Critical for YMYL topics (health, finance, legal).
- Behavioral — what users do in the SERP (do they bounce back, how long they stay).
- Technical — speed (Core Web Vitals), mobile UX, security.
- Local & personal — geolocation, search history, device type.
What you actually see on Google in 2026
The "10 blue links" era is long gone. A modern SERP is a UI assembled live for each query:
| Feature | When it appears |
|---|---|
| 🤖 AI Overviews | AI-generated summary at the top of the SERP for complex/informational queries |
| ⚡ Featured snippets | Direct answer extracted from a page (the "position zero") |
| 📚 Knowledge panels | Entity card on the right: brand, person, place |
| 📍 Map pack (3-pack) | Local intent: 3 businesses on a map |
| ❓ People Also Ask | Related questions with expandable answers |
| 🎬 Video carousel | When video beats text ("how to..." queries) |
| 🛒 Shopping | Transactional queries: products with prices and ratings |
Many users never even reach the classic organic links — the answer is already visible up top. This is called zero-click search. Today, landing in a featured snippet or in the sources cited by AI Overviews is its own SEO objective, every bit as important as "rank in the top 10."
What this means for your site
Each of the 4 stages is a place where you can help Google — or accidentally block it:
| Stage | How to break it | How to help it |
|---|---|---|
| Crawl | Block everything in robots.txt; bury pages deep; break internal links | Sitemap.xml + clean internal linking + clear URL structure |
| Render | Content only appears after click/scroll; critical JS fails to load | SSR/SSG; key content in the initial HTML |
| Index | Accidental noindex; empty title; duplicates without canonical | Unique meaningful content, schema.org, clean meta tags |
| Rank | Weak topical relevance; no links or mentions | Deep useful content; domain authority; healthy Core Web Vitals |
Practice: walk one page through all four stages
Take any page of your site and check it in order. Five steps:
- Open the URL Inspection tool in Search Console and paste the address. "URL is on Google" means the first three stages are done.
- Run a live test and open the HTML tab. That is what the crawler saw. If your copy is missing there, the problem is rendering.
- Look at "Google-selected canonical". If it differs from yours, Google treats the page as a duplicate of another one.
- Check the last crawl date. An old date on an important page means the crawler rarely comes by.
- Walk the whole site with the site crawler, and take the questionable URLs to the indexing audit.
The glossary covers the terms behind each stage: crawlability, JavaScript SEO, indexability, XML sitemap.
How to tell a page cleared every stage
Write down four facts about the page you are checking: its status in URL Inspection, the last crawl date, the Google-selected canonical, and whether it shows up in the page indexing report. That is your "before".
After your fixes, look at the same four. Progress looks like this: the page moves to "URL is on Google", the crawl date refreshes, the canonical matches yours. Rankings are beside the point here — this level only checks admission to the index.
Nobody can promise a timeline: a new page may wait weeks for a crawl, and resubmitting it for indexing does not move the queue.
What comes next
We have covered how Search finds and processes pages. The next lesson is search intent: why two pages with the same words end up in different places. If you skipped the first lesson, start with what SEO is.