Skip to content
🚀
SEO Foundations
Lesson 9 of 16 · AI & Search
FREE +50 XP

How LLMs Work and What They Read on Your Site

LLMs work on a site differently from a person: they do not read a page straight through. They cut the text into pieces, look for the ones that match in meaning, and assemble an answer from what they found. Understanding that mechanism explains why a short self-contained section gets cited more often than a long article — and why "keyword density" does nothing here at all.

Model mechanics are not Google's territory and there is no documentation for them. There is documentation for what follows from them for your site: AI features in Search and how Search works. Those are what the conclusions lean on.

🧑‍💻
Alex · Tuesday, 4:00 PM
"Kate asked: 'How does ChatGPT actually work? Does it read our site?' Alex put down the mouse — the question was good, but explaining took him 30 minutes."

The previous lesson covered where AI answers appear and what Google requires to be in them. This one covers how the machine that assembles the answer actually works. Not for the vocabulary: it is what lets you tell a sensible piece of advice about structure from someone selling "secret markup".

A model does not "understand" — it predicts the next token

This is the most important sentence in the lesson. A large language model is not a mind and not "artificial intelligence" in the science-fiction sense. It is a neural network trained on an enormous corpus of text that does one thing: predicts the next token — a word or part of one. When you type "weather in Kyiv today", the model does not know the weather. It picks the most likely continuation of that phrase.

The sizes of current models are not published by their makers, so the "so many trillion parameters" figures that circulate cannot be checked. For the work it does not matter: three properties carry the practical weight, not the size.

PropertyWhat follows from it for your site
Training stopped on a particular dateRecent information — prices, news, updates — arrives not from training but from a web search at query time. That is your page's way in
A model can invent confidentlyNobody has measured the rate of those errors in a way anyone can reproduce. But they are why answers started citing sources: a link works as an anchor
A model reproduces what is frequentThe more often you are written about in context, the better your odds of appearing in an answer. Mentions work even without a link

The training date: why a model does not know yesterday's news

Every model has a point past which it knows nothing: the date its training data collection ended. Ask about something later than that and, without web access, the model will either say so or invent an answer.

Specific dates per model are deliberately not listed here: they change with every release, and any such list goes stale faster than the lesson. Check the primary source — the model maker's documentation — or simply ask the model and verify against its card.

The practical conclusion is single and stable: for a model to cite you in an answer to a current query, your page has to be found at query time. An individual business cannot train someone else's model on its content. It can get into the search results the model assembles the answer from.

This also answers the common question of whether to block AI bots. Google's user agents and what each one does are listed in the crawler documentation. Blocking them also blocks your way into the answers — a sensible move only when that is genuinely what you want.

Retrieval: how AI search finds your page

Every modern AI answer follows one scheme: search first, generate second. The model does not hold your site in memory — it receives the retrieved fragments as context and writes the answer from them. Hence four steps:

❓
Query
The user asks
→
🔍
Retrieve
The system finds pages
→
✂️
Chunk
Cuts them into pieces
→
💬
Generate
The model writes the answer

At Retrieve the system finds relevant pages in the index. At Chunk it takes not the whole page but the fragment that best answers the question. At Generate the model gets those fragments and assembles an answer. The links under the answer are the addresses the fragments came from.

Classic Search works similarly: Google describes answers resting on a single passage of a page, and for AI features, a query fanning out into several related searches across subtopics. The details are in the section on AI features.

📌 The main implication for content: what gets cited is a fragment, not a page. If you have a superb 5,000-word page with no self-contained piece that answers a specific question on its own, there is nothing to take. The text that wins is built from short dense sections, each closing its own subquestion.

Tokens and embeddings: how a machine compares meaning

Two concepts that explain why "keyword density" stopped working while covering a topic properly still does.

ConceptWhat it isWhat it changes in your text
TokenThe smallest unit of text for a model: a whole word or part of one. Long and rare words split into several tokensThe context window is measured in tokens — how much the model can read at once. Filler spends it for nothing
EmbeddingA fragment of text becomes a set of numbers — coordinates in a space of meaning. "Buy sneakers" and "purchase running shoes" land close together despite sharing no wordsYou do not need to repeat the exact phrasing: closeness is computed on meaning, not on string matches
Vector closenessA measure of how far two sets of coordinates point the same way. It decides which fragment is nearest the queryA dense, specific paragraph gives distinct coordinates; a vague one gives vague coordinates and loses

The practical meaning cuts both ways. Good news: no need to repeat "buy sneakers" seven times — synonyms and context count. Bad news: text made of generalities loses to specific text even when it formally contains the right words. The density analysis and the text quality check show your own page from that angle.

Fragments: how section length affects citation

Search systems cut text into fragments. The size is always a trade-off: a short fragment hits a narrow question precisely but loses context; a long one keeps context but carries extra material and matches a specific query less well. Neither Google nor the assistant makers publish exact sizes, so aiming at "the right number of tokens" is pointless.

What can be said with confidence: a section that reads on its own is easier to cite than a piece torn from the middle of continuous prose. Hence a practical rule that is about meaning rather than counts: each section should answer one question completely, without requiring the neighbours.

What you writeHow a machine reads it
A short section under its own subheading, one question and one answerA ready, self-contained fragment
A long section where the answer accumulates across four paragraphsAny single piece is incomplete, and risky to take
A paragraph opening with "as we said above"Unreadable outside the page entirely

The last row is the most common mistake. In-text references to "above" and "below" tie a paragraph to its page and make it unusable as a quote.

What a page that is easy to cite looks like

AspectAn ordinary pageA page that is easy to quote
Intro"In this article we will tell you… founded in 2010…"A direct answer in the first two or three sentences
Headings"Introduction", "Overview", "Conclusion"The questions themselves: "How much does X cost", "How to set up Y"
Section lengthFifteen hundred words under one subheadingA section that reads on its own and closes one subquestion
Claims"This is a very important factor" — with nothing behind it"Google shows FAQ results only for government and health sites" — with the announcement linked and dated
SentencesLong, with three subordinate clauses and caveatsShort: "X is Y, because Z"
AuthorAnonymous, or "our team of experts"A name, what they do, a link to their page

On that last row: the author name and the update date can be mirrored in markup — the fields are described in the Article schema. That is not "markup for AI", which does not exist, but ordinary structured data; the term is covered in the glossary: structured data.

Six rules for text that is easy to cite

1. One section, one question

Reading only that section, a person should grasp its topic without the rest of the page. That is the self-containment everything else is for. If a section has grown and now answers three questions, split it with subheadings.

2. Short sentences instead of subordinate clauses

"SEO is the work of getting a page found in search" reads better than "SEO, understood broadly as the process of optimizing websites for search engines with a view to improving their visibility in organic results". The second smears the meaning and matches a query worse — and is harder for a human to read.

3. Headings are questions

Not "Benefits of X" but "Why you need X". Not "Technical specifications" but "How much does X weigh". The heading sets the fragment's topic, and the closer it is to how people ask, the easier the fragment is to find. What phrasing your audience uses is visible in the AI overview analyzer.

4. Numbers with a source and a date — and no others

Back any claim you can with a checkable number. Checkable specifically: a link, a date, who measured it. Invented statistics with a source that does not exist — and this topic is full of them — work against you the moment a reader decides to check.

5. Write the FAQ for people, not for the markup

A block of questions at the end is useful: it naturally splits a topic into short self-contained answers. But FAQPage markup no longer earns a rich result: in August 2023 Google limited those to well-known authoritative government and health sites. You can leave the markup in place, it does no harm, but do not count on it as a "citation accelerator".

6. Do not tie paragraphs to the page

"As we wrote above", "in the previous section", "see the table below" — each of these makes the fragment unreadable out of context. If you need a reference, make it explicit: not "above" but "in the section on tokens", or an ordinary link.

⚡ A two-minute page check: open your main article and walk through five points. Is there a direct answer to the main question in the first two or three sentences? Is every subheading a question? Is there at least one section you could quote whole and have it still make sense? Are there numbers with a source and a date? Is the author named? Four yeses out of five and the page can be taken into an answer. One or two, and it needs rewriting.

Practice: break your page into fragments

Four steps on your own text:

  1. Take the page with the most impressions in the performance report and list its subheadings.
  2. Next to each, write the question that section answers. If no question forms, or three do, the section needs splitting or rewriting.
  3. Read each section alone, as if the neighbours were gone. Mark every place the text leans on "above", "below" and "as already mentioned" — that is the tie to the page.
  4. Check readability with the readability analysis, and topic coverage with entity coverage: it shows the subquestions the page does not have.

The glossary covers the terms: AI Overviews, E-E-A-T, featured snippet, structured data.

How to tell the page became quotable

There is one checkable sign, and it has nothing to do with AI: take any section of your page and read it aloud to someone who has not opened the article. If they got the point whole, the fragment is self-contained. If they asked "what is this about?", it is not.

The second sign is in the data. Compare the page's impressions and clicks before and after the rewrite over equal windows. Impressions growing at the same positions means the page now matches more phrasings, which is exactly what splitting into subquestions is for.

What not to expect: that structure replaces substance. A beautifully sliced text about nothing has as little to quote as an unbroken wall of words. Structure helps find the point; it does not create one.

What comes next

The next lesson is AI in content creation: where the machine can be trusted with the work and where it costs you positions. The previous lesson, on AI answers in the results, is here.

🧑‍💻
Kate
"So AI doesn't 'read' like a human. It cuts content into pieces, finds similar pieces, and assembles an answer from them?"
😌
Alex
"Exactly. So good writing is not 'a long article' — it is several short self-contained sections assembled into one coherent piece. Every section is a potential citation. What adds weight to it is a checkable number with a source and the name of someone who knows the subject."
🎮 Test yourself: answer the question in the task!
🎯
Lesson Task
Test your knowledge and earn +20 XP
← AI Overviews and AI Mode in Search
Lesson 9 of 16
Go to Task →