AEO Metrics: How to Track Your Brand's AI Visibility
The metrics that tell you whether AI engines recommend you, how each one is computed, the benchmarks we actually see, and the mistakes that make a dashboard lie.
What gets measured gets improved — if it is measured honestly
Tracking AI visibility needs different metrics from SEO. There are no rankings to read off a results page: an engine writes an answer, and the question is whether you are in it, where, and on whose evidence. Below is the measurement framework AEOTrack uses, with the benchmarks from the brands it tracks, and the four ways a number goes wrong.
The five core metrics
1. Mention rate
Formula: answers that name your brand ÷ usable answers, per engine, on non-branded questions only.
Two qualifiers do the work. Usable excludes answers where the engine errored or was cut off before it reached the point of naming anyone; scoring those as misses punishes you for the provider's bad day. Non-branded excludes questions that contain your name or a near-spelling of it, which any engine will surface for lexical reasons. Benchmarks on our data: most brands start under 10%; 10–25% is an established brand in its niche; above 25% on non-branded questions is category leadership.
2. Citation rate
Formula: answers where the engine cites a page on your domain as a source ÷ usable answers.
Citation and mention are different events and fail for different reasons. Across 446 answers for one tracked domain, 68 cited the site and 38 named it, and on one engine 22 answers cited it without naming it once. A brand that is cited but not mentioned is being used as evidence for someone else's recommendation. Report the two numbers separately; a blended "visibility" figure hides the gap.
3. Position in answer
Formula: when the engine ranks options, your rank, discounted logarithmically (1 ÷ log₂(rank + 1)) so that #1 against #3 matters far more than #8 against #10.
Being named last in a list of eight and being named first are both "mentioned". They are not the same outcome for a buyer reading the answer. Position carries the largest weight in AEOTrack's composite because it is the thing a customer can most directly move — and it is the term that was pinned at a constant in most trackers until parsers learned to read the bold-heading lists engines actually write.
4. Share of voice
Formula: your mentions ÷ total mentions of every brand in your named competitor set, over the same answers.
Share of voice only means something against a fixed set. Adding a competitor pushes everyone's share down without anything about your visibility changing, so choose three to five, hold the set for a quarter, and note any change in the report. Pair it with a per-engine breakdown: a rival can dominate one engine and be absent from another.
5. Sentiment
Formula: a −100 to +100 rating of how favourably each engine's answer describes you, scored per answer rather than once per question.
The least-weighted term, and the one to watch on branded questions specifically: "is X any good" is where an engine repeats what review sites say about you.
The composite, and the confidence range
AEOTrack's score is 25 points for being present at all plus up to 75 for how you are present (45% position, 30% citation, 25% sentiment), averaged across engines by usage weight and across questions with a 30-day half-life so last quarter's answers fade. Every score carries a confidence range. AI answers vary between runs; a score without a range moves on its own and teaches you to ignore it. Alerts fire only when the new range no longer overlaps the old one.
How to track these metrics
By hand
A spreadsheet with one row per question per engine per run: date, engine, named (yes/no), position, cited (yes/no), competitors named, sources cited. It works for ten questions on two engines for about a month, and then it stops, because AI answers are non-deterministic and a single run is an anecdote rather than a rate.
Automated
| Capability | Why it matters |
|---|---|
| Scheduled runs | A rate needs repetition; a trend needs a schedule |
| Live search on every engine | Model memory is not what a user sees |
| Per-engine reporting | Engines disagree more than they agree |
| Stored answers and citations | The "why" lives in the answer text and the source list |
| Confidence ranges | Tells a real change from run-to-run noise |
| Branded / near-brand detection | Stops a category-named brand from flattering itself |
Four ways the number lies
- Counting branded questions. A brand that tracked eleven questions, five of which contained its category name, scored 17; on neutral questions alone it scored 10.
- Scoring cut-off answers as misses. Reasoning models spend output tokens thinking; cap them too low and 90% of an engine's answers stop before naming anyone.
- Substring matching. "AEO Track" matching inside "AEO tracking", or a competitor called Profound being credited for "a profound impact". The letters of a common word are not a brand.
- Blending engines. One number over five engines hides which one is failing, and the fixes differ by engine.
A weekly cadence that holds
Scheduled checks run weekly on every plan. Monday, read the per-engine table and the alerts that cleared the noise band. Tuesday, open the sources cited for the questions where a competitor was named and you were not. Wednesday to Friday, work the top three of those sources. Do not re-run checks daily to watch the number; the answers change over weeks, and daily re-runs measure the engine's variance, not your progress.
See how your site performs in AI answers
Start your 14-day free trial — every feature unlocked, no credit card required.
Start Free Trial →