Blog/Plain-English Guide
PLAIN-ENGLISH GUIDE

How Do AI Overviews Choose Their Sources? What Gets a Page Into the Box

Six tall dark blank upright panels standing in a row on a low rail in a deep blue void, the second and fifth panels pulled forward and standing proud of the row with the empty channels behind them washed warm amber

Key Takeaways

  • Stop thinking of it as a ranking and start thinking of it as an assembly. A box is built from several pages at once, each one supplying a different part of the answer. Nobody wins the whole thing, so being second is not the reason you are absent.
  • Be in the index first, because eligibility is a gate and not a score. Google says the sources it links come from its regular Search systems. A page that was never fetched and kept cannot be considered, however good it is.
  • Write for the narrower question, not the headline one. One typed question is broken into several smaller ones. Each smaller one goes to whichever page contains that specific thing, which is usually a price, a timeframe or a named situation.
  • Put the liftable sentence where a stranger can find it in five seconds. A passage that answers a question completely, in one or two sentences, near a heading that asks it, is the unit that gets quoted. Buried context does not get lifted.
  • Expect corroboration to matter more than polish. When two sources agree on a fact, a model repeats it with more confidence. One page saying something no other page says is the hardest thing to get quoted.

You searched something your buyers search, and Google answered it in a box before the results. Four sites were linked beside it. You were not one of them, even though two of the four sit below you in the normal results.

That combination confuses people, and it should. It looks like a ranking and behaves like something else.

This post covers what the box is actually doing when it picks sources, why position and citation come apart, and the two or three things that decide whether a page of yours can be used at all.

How do AI Overviews choose their sources?

They break your question into smaller ones, then pull a different page in to answer each part. The pages come from what Google has already indexed.

Google's own documentation on AI features in Search is the plainest statement of the rule: there is no separate submission process and no special markup that qualifies a page, and the links shown are drawn from the same Search systems that produce ordinary results. Two things follow from that. Being in the index is a gate you pass or fail. And once you are through it, what decides the outcome is what your page actually says.

So the honest answer is a sequence. Be findable. Contain the specific thing the smaller question is asking for. Say it in a way a stranger could repeat without reading the rest of the page.

Why this is an assembly and not a ranking

Because no single page is chosen as the winner. Several are used at once, and each is used for a different sentence.

Think about how the box is written rather than how it looks. It opens with a definition, then gives a range, then adds a caveat, then names a couple of examples. Those four moves do not have to come from one place, and usually do not. A page that carries one of them well is more useful to the box than a page that covers all four vaguely.

This is why the ranking comparison misleads. You are not losing a contest to the pages that got cited. You did not have the part they were carrying.

ONE QUESTION, FOUR SMALLER ONESILLUSTRATIVE
What the buyer typed Who is the best coach for someone like me?
What does it cost? A page with a number on it A competitor with a pricing page
Who does this for my situation? A page naming the situation A specialty page
How long does it take? A stated timeframe An FAQ answer
Is it any good? Something a third party said A review profile

Four different pages answered one question. The fourth part was carried by a page nobody wrote: a review profile. An illustrative example, not measured data.

answerhalo.com
Each smaller question goes to whichever page holds that specific thing. Winning the headline question is not what is being decided.

It also explains a pattern that feels unfair. A thin page with a clear price gets cited over a thorough page with no price, because the smaller question was about cost and only one of them answered it. The thorough page was better and still had nothing to lift.

One question becomes several questions

Because most real questions contain more than one question, and the box has to satisfy all of them before it is worth showing.

Read your own service page against the four parts in the figure above. Most pages answer the second one, sometimes the third, and skip the first entirely. Price is the single most common gap, and it is the part buyers ask about first. The same pattern shows up in what kinds of pages AI assistants quote most, where the pages that get lifted are rarely the ones the owner is proudest of.

This is not unique to Google, either. The mechanism differs but the consequence does not, which is worth reading next to how Perplexity chooses its sources. Different plumbing, same requirement: contain the specific thing.

What a page has to contain to get lifted

A self-contained passage that answers a question completely, sitting near a heading that asks it, in a page that was indexed months ago rather than last night.

Google's guidance on creating helpful content describes the standard without ever mentioning AI: a page written for a person with the question, answering it well enough that they leave satisfied. That standard is doing the work here, because a passage that satisfies a reader is also the passage that survives being pulled out of its page.

There is one research finding worth knowing alongside it. A Princeton-led study, GEO: Generative Engine Optimization, tested content changes against generative search answers and found that adding citations, quotations and statistics raised how visible a source was in the generated answer, by roughly 30 to 40 percent on their benchmark. It is one academic benchmark rather than Google's live system, so treat the direction as the finding and not the number.

The four parts of one buyer question, what each part needs from a page, and what usually fails to supply it
The smaller question What a page must contain What usually supplies it Why most sites miss it
What does it cost? An actual number, or a stated range A pricing page Price is held back for the call
Who does this for my situation? The situation named in words buyers use A specialty page One page tries to serve everybody
How long does it take? A timeframe stated as a sentence An FAQ answer The timeframe is only said out loud
Is it any good? Something a third party wrote A review profile or a listing Praise lives on the site itself

Read the last row carefully, because it is the one you cannot solve by writing. Nothing you publish about yourself supplies it. That part gets carried by pages you do not own, which is the whole argument for making sure those pages exist and say something specific.

What to do about it this month

Pick the one question your buyers ask before they contact you. Then check whether any page you own answers it in a single passage, near a heading, in words a stranger would repeat. Most owners find the answer exists somewhere on the site, in the middle of a paragraph about something else.

Move it. Give it a heading phrased as the question, put the answer in the first two sentences under it, and keep it self-contained enough to make sense with the rest of the page removed. That is the whole technique.

Then measure instead of guessing, because asking once proves nothing. The method is set out in how to check if AI recommends you, and the background on the box itself is in what an AI Overview on Google actually is.

If you would rather have the first reading done for you, the free check asks ChatGPT 25 real buyer questions from your niche at $0 and emails a plain-English readout: your score, who got named in your place, and the one fix to start with. It runs as soon as you ask for it, and the report lands within the hour between 7am and 5pm Central, or by 10am the next morning outside those hours.

Questions we hear the most

How do AI Overviews choose their sources?

They assemble an answer from pages already in Google Search, picking a different page for each part of the question. Google says the links shown come from its regular Search systems, so being indexed is the gate, and containing the specific thing being asked decides the rest.

Is the number one result always cited in the AI Overview?

No. The box often cites pages that do not hold the top spot, and often skips pages that do. The two things are decided differently, which is why the overlap between them is partial rather than complete.

Do I need special markup to be cited?

No. Structured data helps machines understand a page, but there is no schema type that makes a page eligible for an AI Overview. Being findable and containing a clear, quotable answer does far more than markup does.

Why does the box cite four sites I out-rank?

Because each of them supplied a different piece. One had a price, one had a timeframe, one had a specialty statement. Out-ranking a page on one search does not mean you carry the fact that page was quoted for.

Can I stop my pages appearing in AI Overviews?

Partly. Google offers snippet controls that limit how much of a page can be shown, and blocking a page from Search removes it from these features too. Most small businesses want the opposite problem solved.

How long does it take for a new page to be considered?

There is no published timeline. The page has to be crawled and kept first, and only then can it be reached for an answer, so treat the first few weeks as an observation window rather than a verdict.

How do I find out whether AI is naming me right now?

The AnswerHalo free check asks ChatGPT 25 real buyer questions from your niche and emails a plain-English report: your score, who got named in your place, and the one fix to start with. It costs $0 and it runs as soon as you ask for it.

What does the paid audit add that the free check does not?

The $500 audit runs 125 checks across five AI tools instead of 25 on one, records what got named and what got cited on each, and ends with the page that wins your worst question, specified section by section. It is in your inbox by this time tomorrow.

SEE WHERE YOU STAND

Your buyers are already asking. Find out what AI tells them.

The free check asks ChatGPT 25 real buyer questions about your niche and sends you the report within the hour, at $0. If it shows you are already getting named everywhere, we will say so in plain words.

Get my free check →