How LLMs Work and What They Read on Your Site
LLMs work on a site differently from a person: they do not read a page straight through. They cut the text into pieces, look for the ones that match in meaning, and assemble an answer from what they found. Understanding that mechanism explains why a short self-contained section gets cited more often than a long article — and why "keyword density" does nothing here at all.
Model mechanics are not Google's territory and there is no documentation for them. There is documentation for what follows from them for your site: AI features in Search and how Search works. Those are what the conclusions lean on.
The previous lesson covered where AI answers appear and what Google requires to be in them. This one covers how the machine that assembles the answer actually works. Not for the vocabulary: it is what lets you tell a sensible piece of advice about structure from someone selling "secret markup".
A model does not "understand" — it predicts the next token
This is the most important sentence in the lesson. A large language model is not a mind and not "artificial intelligence" in the science-fiction sense. It is a neural network trained on an enormous corpus of text that does one thing: predicts the next token — a word or part of one. When you type "weather in Kyiv today", the model does not know the weather. It picks the most likely continuation of that phrase.
The sizes of current models are not published by their makers, so the "so many trillion parameters" figures that circulate cannot be checked. For the work it does not matter: three properties carry the practical weight, not the size.
| Property | What follows from it for your site |
|---|---|
| Training stopped on a particular date | Recent information — prices, news, updates — arrives not from training but from a web search at query time. That is your page's way in |
| A model can invent confidently | Nobody has measured the rate of those errors in a way anyone can reproduce. But they are why answers started citing sources: a link works as an anchor |
| A model reproduces what is frequent | The more often you are written about in context, the better your odds of appearing in an answer. Mentions work even without a link |
The training date: why a model does not know yesterday's news
Every model has a point past which it knows nothing: the date its training data collection ended. Ask about something later than that and, without web access, the model will either say so or invent an answer.
Specific dates per model are deliberately not listed here: they change with every release, and any such list goes stale faster than the lesson. Check the primary source — the model maker's documentation — or simply ask the model and verify against its card.
The practical conclusion is single and stable: for a model to cite you in an answer to a current query, your page has to be found at query time. An individual business cannot train someone else's model on its content. It can get into the search results the model assembles the answer from.
This also answers the common question of whether to block AI bots. Google's user agents and what each one does are listed in the crawler documentation. Blocking them also blocks your way into the answers — a sensible move only when that is genuinely what you want.
Retrieval: how AI search finds your page
Every modern AI answer follows one scheme: search first, generate second. The model does not hold your site in memory — it receives the retrieved fragments as context and writes the answer from them. Hence four steps:
At Retrieve the system finds relevant pages in the index. At Chunk it takes not the whole page but the fragment that best answers the question. At Generate the model gets those fragments and assembles an answer. The links under the answer are the addresses the fragments came from.
Classic Search works similarly: Google describes answers resting on a single passage of a page, and for AI features, a query fanning out into several related searches across subtopics. The details are in the section on AI features.
Tokens and embeddings: how a machine compares meaning
Two concepts that explain why "keyword density" stopped working while covering a topic properly still does.
| Concept | What it is | What it changes in your text |
|---|---|---|
| Token | The smallest unit of text for a model: a whole word or part of one. Long and rare words split into several tokens | The context window is measured in tokens — how much the model can read at once. Filler spends it for nothing |
| Embedding | A fragment of text becomes a set of numbers — coordinates in a space of meaning. "Buy sneakers" and "purchase running shoes" land close together despite sharing no words | You do not need to repeat the exact phrasing: closeness is computed on meaning, not on string matches |
| Vector closeness | A measure of how far two sets of coordinates point the same way. It decides which fragment is nearest the query | A dense, specific paragraph gives distinct coordinates; a vague one gives vague coordinates and loses |
The practical meaning cuts both ways. Good news: no need to repeat "buy sneakers" seven times — synonyms and context count. Bad news: text made of generalities loses to specific text even when it formally contains the right words. The density analysis and the text quality check show your own page from that angle.
Fragments: how section length affects citation
Search systems cut text into fragments. The size is always a trade-off: a short fragment hits a narrow question precisely but loses context; a long one keeps context but carries extra material and matches a specific query less well. Neither Google nor the assistant makers publish exact sizes, so aiming at "the right number of tokens" is pointless.
What can be said with confidence: a section that reads on its own is easier to cite than a piece torn from the middle of continuous prose. Hence a practical rule that is about meaning rather than counts: each section should answer one question completely, without requiring the neighbours.
| What you write | How a machine reads it |
|---|---|
| A short section under its own subheading, one question and one answer | A ready, self-contained fragment |
| A long section where the answer accumulates across four paragraphs | Any single piece is incomplete, and risky to take |
| A paragraph opening with "as we said above" | Unreadable outside the page entirely |
The last row is the most common mistake. In-text references to "above" and "below" tie a paragraph to its page and make it unusable as a quote.
What a page that is easy to cite looks like
| Aspect | An ordinary page | A page that is easy to quote |
|---|---|---|
| Intro | "In this article we will tell you… founded in 2010…" | A direct answer in the first two or three sentences |
| Headings | "Introduction", "Overview", "Conclusion" | The questions themselves: "How much does X cost", "How to set up Y" |
| Section length | Fifteen hundred words under one subheading | A section that reads on its own and closes one subquestion |
| Claims | "This is a very important factor" — with nothing behind it | "Google shows FAQ results only for government and health sites" — with the announcement linked and dated |
| Sentences | Long, with three subordinate clauses and caveats | Short: "X is Y, because Z" |
| Author | Anonymous, or "our team of experts" | A name, what they do, a link to their page |
On that last row: the author name and the update date can be mirrored in markup — the fields are described in the Article schema. That is not "markup for AI", which does not exist, but ordinary structured data; the term is covered in the glossary: structured data.
Six rules for text that is easy to cite
1. One section, one question
Reading only that section, a person should grasp its topic without the rest of the page. That is the self-containment everything else is for. If a section has grown and now answers three questions, split it with subheadings.
2. Short sentences instead of subordinate clauses
"SEO is the work of getting a page found in search" reads better than "SEO, understood broadly as the process of optimizing websites for search engines with a view to improving their visibility in organic results". The second smears the meaning and matches a query worse — and is harder for a human to read.
3. Headings are questions
Not "Benefits of X" but "Why you need X". Not "Technical specifications" but "How much does X weigh". The heading sets the fragment's topic, and the closer it is to how people ask, the easier the fragment is to find. What phrasing your audience uses is visible in the AI overview analyzer.
4. Numbers with a source and a date — and no others
Back any claim you can with a checkable number. Checkable specifically: a link, a date, who measured it. Invented statistics with a source that does not exist — and this topic is full of them — work against you the moment a reader decides to check.
5. Write the FAQ for people, not for the markup
A block of questions at the end is useful: it naturally splits a topic into short self-contained answers. But FAQPage markup no longer earns a rich result: in August 2023 Google limited those to well-known authoritative government and health sites. You can leave the markup in place, it does no harm, but do not count on it as a "citation accelerator".
6. Do not tie paragraphs to the page
"As we wrote above", "in the previous section", "see the table below" — each of these makes the fragment unreadable out of context. If you need a reference, make it explicit: not "above" but "in the section on tokens", or an ordinary link.
Practice: break your page into fragments
Four steps on your own text:
- Take the page with the most impressions in the performance report and list its subheadings.
- Next to each, write the question that section answers. If no question forms, or three do, the section needs splitting or rewriting.
- Read each section alone, as if the neighbours were gone. Mark every place the text leans on "above", "below" and "as already mentioned" — that is the tie to the page.
- Check readability with the readability analysis, and topic coverage with entity coverage: it shows the subquestions the page does not have.
The glossary covers the terms: AI Overviews, E-E-A-T, featured snippet, structured data.
How to tell the page became quotable
There is one checkable sign, and it has nothing to do with AI: take any section of your page and read it aloud to someone who has not opened the article. If they got the point whole, the fragment is self-contained. If they asked "what is this about?", it is not.
The second sign is in the data. Compare the page's impressions and clicks before and after the rewrite over equal windows. Impressions growing at the same positions means the page now matches more phrasings, which is exactly what splitting into subquestions is for.
What not to expect: that structure replaces substance. A beautifully sliced text about nothing has as little to quote as an unbroken wall of words. Structure helps find the point; it does not create one.
What comes next
The next lesson is AI in content creation: where the machine can be trusted with the work and where it costs you positions. The previous lesson, on AI answers in the results, is here.