Skip to content
🚀
SEO Foundations
Lesson 2 of 16 · Introduction to SEO
FREE +50 XP

How Search Engines Work

Search engines work in four stages: crawling, rendering, indexing and ranking. Until a page has cleared the first three it cannot appear for any query at all, however well the copy is written. This lesson walks through each stage, where pages usually get stuck, and how to check your own.

Google describes the overall flow in «How Google Search works» and lists its robots in the crawler overview.

🧑‍💻
Alex · Monday, 3:30 PM
"Kate, the new intern, walked over with her laptop: 'Alex, why isn't our site showing up on Google? We launched it three weeks ago...'"

Simple answer: the search engine doesn't know your site exists yet. When you search Google, the answer is already prepared — Google has crawled the web in advance, read billions of pages, organized them, and now just picks the relevant ones. Let's break down what's happening behind the scenes.

A search engine isn't an AI chat. It's a giant pre-built database

When you type a query, Google doesn't run across the web in real time — that would be catastrophically slow. Instead, it gathers information about every page in advance into a massive database called the index. Your query then hits the index in milliseconds.

The pipeline has four sequential stages. At any one of them your site can get stuck — and then no amount of SEO will help.

🕷️
Crawl
Bot finds the page
→
🎬
Render
Runs it like a browser
→
📚
Index
Stores it on the shelves
→
🏆
Rank
Picks answers per query

Stage 1. Crawling — discovering pages

Googlebot is a "spider" program. It finds new URLs in two ways:

Which URLs the crawler visits and which it skips comes down to your sitemap and robots.txt: how Google reads a sitemap and what robots.txt does.

  • By following links — it hops from a known page to anything it links to. If any indexed site links to yours, the bot eventually arrives.
  • Via sitemap.xml — if you've submitted a sitemap in Google Search Console, the bot gets your URL list directly.
⚠️ 4 reasons Googlebot may not reach a page:
  1. Nothing links to it and it's not in the sitemap — the bot has no idea it exists
  2. It's buried too deep (5+ clicks from the homepage) — the bot decides it's not worth the cost
  3. The bot crawled it before and decided it's not worth a recrawl (thin content, duplicate)
  4. Your crawl budget is exhausted — the cap on how many pages Google is willing to fetch from your domain per day

Crawl budget is a real constraint. The web is too big for Google to crawl exhaustively, so it prioritizes: large authoritative sites get crawled often, small new ones much less frequently.

Stage 2. Rendering — Google actually executes the page

This step didn't exist a decade ago. Back then Google just read the raw HTML. In 2026, most sites are JavaScript-driven — React, Vue, Next.js — and the actual content only shows up after scripts run.

What Google does with JavaScript, and why content drawn by a script reaches the index later, is covered in the JavaScript SEO basics.

So Google now runs the page like a real browser (Chromium): loads HTML, executes JavaScript, waits for things to settle, and only then "snapshots" the final DOM. This step is called rendering.

💡 Practical takeaway: if your content only appears after a click, scroll, or AJAX call, there's a real risk Googlebot won't see it. Fixes: server-side rendering (SSR), static generation (SSG), or plain HTML enhanced progressively with JS.

Stage 3. Indexing — putting things on shelves

Google semantically analyzes the rendered page: it parses headings, text, images, structure, meta tags, and schema.org markup, extracts key concepts, and stores them in the index.

This is where Google picks a canonical URL among similar pages; the rules are in the canonicalization documentation.

The modern index is not just a text database. It's a connected web of stores:

StoreWhat it holds
Main indexMetadata for every page: text, headings, links, sitemap data
Vertical indexesSeparate stores for images, videos, news, products — each with its own attributes
Knowledge GraphA graph of entities: people, companies, places, concepts, and the relationships between them. Sources include Wikipedia, Wikidata, and schema markup

The Knowledge Graph is why Google understands that "F1" = Formula One = motor racing, that "Elon Musk" is connected to Tesla and SpaceX, that "buy" and "sell" are close concepts. Modern search isn't word matching — it's entity matching.

⚠️ Being in the index ≠ being in the top results. The index holds billions of pages. Page 1 of the SERP holds ten. Between them sits the ranking step.

Stage 4. Ranking — picking the best answers

When a user submits a query, Google in milliseconds:

Google lists the systems involved in ranking in its guide to ranking systems.

  1. Identifies the query's intent (informational / transactional / local)
  2. Pulls relevant pages from the index
  3. Ranks them by hundreds of signals and assembles the SERP

In 2026, a major share of ranking is driven by machine learning. Google's key ML systems:

SystemWhat it does
RankBrain (2015)Understands the meaning of rare or long-tail queries via embeddings
BERT (2019)Reads each word in context of the full query (word order matters: "ticket to Boston from NYC" ≠ "ticket from Boston to NYC")
MUM (2021)Connects topics across languages and content types (text + video + images together)
SpamBrainAuto-detects and demotes spam, thin content, and low-quality AI-generated junk

The main ranking signal groups:

  • Relevance — does the page match the query. The largest factor by far.
  • Authority — who links to you, how often you're cited in your niche (the heir to PageRank).
  • EEAT — Experience, Expertise, Authoritativeness, Trustworthiness. Critical for YMYL topics (health, finance, legal).
  • Behavioral — what users do in the SERP (do they bounce back, how long they stay).
  • Technical — speed (Core Web Vitals), mobile UX, security.
  • Local & personal — geolocation, search history, device type.

What you actually see on Google in 2026

The "10 blue links" era is long gone. A modern SERP is a UI assembled live for each query:

FeatureWhen it appears
🤖 AI OverviewsAI-generated summary at the top of the SERP for complex/informational queries
⚡ Featured snippetsDirect answer extracted from a page (the "position zero")
📚 Knowledge panelsEntity card on the right: brand, person, place
📍 Map pack (3-pack)Local intent: 3 businesses on a map
❓ People Also AskRelated questions with expandable answers
🎬 Video carouselWhen video beats text ("how to..." queries)
🛒 ShoppingTransactional queries: products with prices and ratings

Many users never even reach the classic organic links — the answer is already visible up top. This is called zero-click search. Today, landing in a featured snippet or in the sources cited by AI Overviews is its own SEO objective, every bit as important as "rank in the top 10."

What this means for your site

Each of the 4 stages is a place where you can help Google — or accidentally block it:

StageHow to break itHow to help it
CrawlBlock everything in robots.txt; bury pages deep; break internal linksSitemap.xml + clean internal linking + clear URL structure
RenderContent only appears after click/scroll; critical JS fails to loadSSR/SSG; key content in the initial HTML
IndexAccidental noindex; empty title; duplicates without canonicalUnique meaningful content, schema.org, clean meta tags
RankWeak topical relevance; no links or mentionsDeep useful content; domain authority; healthy Core Web Vitals
🧑‍💻
Alex · 4:15 PM
"Got it now, Kate? Site — three weeks old, no external links, sitemap never submitted. The bot literally hasn't found you yet. It's not 'Google hates us' — it's 'Google has no idea we exist.' Let's set up Search Console, submit the sitemap, and check back in a few days."
🎮 Test yourself: put the search engine stages in the right order in the task!

Practice: walk one page through all four stages

Take any page of your site and check it in order. Five steps:

  1. Open the URL Inspection tool in Search Console and paste the address. "URL is on Google" means the first three stages are done.
  2. Run a live test and open the HTML tab. That is what the crawler saw. If your copy is missing there, the problem is rendering.
  3. Look at "Google-selected canonical". If it differs from yours, Google treats the page as a duplicate of another one.
  4. Check the last crawl date. An old date on an important page means the crawler rarely comes by.
  5. Walk the whole site with the site crawler, and take the questionable URLs to the indexing audit.

The glossary covers the terms behind each stage: crawlability, JavaScript SEO, indexability, XML sitemap.

How to tell a page cleared every stage

Write down four facts about the page you are checking: its status in URL Inspection, the last crawl date, the Google-selected canonical, and whether it shows up in the page indexing report. That is your "before".

After your fixes, look at the same four. Progress looks like this: the page moves to "URL is on Google", the crawl date refreshes, the canonical matches yours. Rankings are beside the point here — this level only checks admission to the index.

Nobody can promise a timeline: a new page may wait weeks for a crawl, and resubmitting it for indexing does not move the queue.

What comes next

We have covered how Search finds and processes pages. The next lesson is search intent: why two pages with the same words end up in different places. If you skipped the first lesson, start with what SEO is.

🎯
Lesson Task
Test your knowledge and earn +20 XP
← What is SEO?
Lesson 2 of 16
Go to Task →