Google AI Overviews vs AI Mode vs ChatGPT: What Each One Cites and Why It Matters
Google now answers questions in two different places, and neither behaves like ChatGPT. If you only track one of the three, you are measuring a third of the market and calling it visibility. This is a plain description of how each surface chooses what to say, what each one tends to cite, and how to track them without fooling yourself.
Three surfaces, one question
| Surface | When it appears | Where it retrieves from | How it cites |
|---|---|---|---|
| Google AI Overviews | Only for queries Google decides need a summary — roughly a third of informational queries, far fewer transactional ones | Google's index, weighted towards pages that already rank on the first page | A short list of linked sources beside the summary; often the same pages that rank organically |
| Google AI Mode | Every query, in the AI Mode tab, with follow-ups | Google's index via a "query fan-out": the question is split into several searches and the results merged | Inline references per claim; more sources than an Overview, from deeper in the rankings |
| ChatGPT (search on) | Whenever the model decides the question needs current information, which for "best X" questions is nearly always | OpenAI's own crawl plus Bing's index; pages that block GPTBot or OAI-SearchBot are invisible | Numbered citations after sentences; typically 3–8 sources per answer |
What that means for what gets cited
AI Overviews reward what already ranks. If you are on page one for a query, you are a candidate for the Overview; if you are not, you almost never are. This is the surface where classic SEO still carries the most weight, and where a brand with strong rankings can be visible in AI without any specific AEO work.
AI Mode rewards being findable for the sub-questions. Because the question is fanned out into several searches, a page that answers one narrow sub-question well — "does tool X integrate with Shopify" — gets pulled into answers to the broad question. Specific, well-titled pages win here over comprehensive guides.
ChatGPT rewards third-party corroboration. Its sources for buyer questions are dominated by roundups, comparison posts, review sites and community threads. Your own site is usually one source among seven or eight, and the recommendation follows the majority. This is the surface where the "get mentioned on the pages the engine already reads" strategy matters most.
How much they disagree
More than anyone expects. On the buyer questions we track, the set of brands named by ChatGPT and the set named by a Google surface for the same question overlap by well under half, and the cited sources overlap by less than a third. A brand can be the first recommendation in one and absent from the other. This is why a single blended "AI visibility" number is misleading: it hides which surface is failing, and the fixes are different for each.
Tracking each one honestly
- Track per surface, never blended. Your dashboard should be able to say "named by ChatGPT in 6 of 10, by Google AI Mode in 1 of 10". The blended figure of 35% would tell you nothing you can act on.
- Use the surface's live retrieval, not the model's memory. A ChatGPT call without search on measures what the model memorised in training, which is not what a user sees. The same is true of Gemini without grounding.
- Separate branded from non-branded questions. All three surfaces will name you when the question contains your name. Only the questions that do not measure discovery.
- Expect variance, and report it. Two runs of the same question on the same surface do not always name the same brands. A score without a confidence range will move on its own and teach you to ignore it.
Where AEOTrack stands
AEOTrack tracks ChatGPT, Gemini, Claude, Perplexity and Grok with live search on every check, and reports each engine separately with a confidence range and a branded/non-branded split. Google AI Mode is being added as a sixth engine, read through the same pipeline so its answers score on the same scale; AI Overviews are next, once a check can be run without paying for two searches to find that no Overview was shown. If your buyers live in Google's surfaces, that is the roadmap to hold us to.
See what the engines say about you.
Presence, position, citations and sentiment — across ChatGPT, Gemini, Claude, Perplexity and Grok. Free for 14 days, no credit card.
Start Your Free Trial →Score what actually matters.
Presence, position, citations and sentiment — across ChatGPT, Gemini, Claude, Perplexity and Grok. Free for 14 days, no credit card.
Start Your Free Trial →