← Back to Blog
Methodology August 12, 2026 8 min read

What an AI Visibility Score Should Actually Measure

Being named first with a citation and being named eleventh in a throwaway list are not the same outcome. Most AI visibility scores can't tell them apart — which means they can't reward the work that matters most.

Most AI visibility scores answer one question: were you mentioned, yes or no? Count the yeses, divide by the total, call it a percentage.

It's simple, it's easy to explain, and it throws away nearly everything that matters.

What a binary score can't see

Consider two AI answers to "best accounting software for freelancers".

Answer A opens with your product, describes it as the standout choice for freelancers, and links to your pricing page as a source.

Answer B lists eleven tools. You're eleventh. The line reads "others include…" and your name appears with no link and no description.

To a binary score these are identical. Both are a mention. Both count as one.

They are obviously not identical. One is a recommendation that will send you customers. The other is a name in a list nobody reads to the end of. And crucially: the work needed to move from B to A is real, valuable work that a binary score will completely fail to register.

The perverse incentive

If your score can't tell position 1 from position 11, it can't reward the improvements that matter most — and it tells you you're done as soon as you scrape into a list.

The four things worth measuring

Presence

Were you named at all? This is the foundation — the binary signal, still necessary, just not sufficient. It earns a baseline share of the score.

Prominence

Where in the answer did you appear? Being first in a ranked list is worth substantially more than being eighth, and the gap between first and third matters more than the gap between eighth and tenth. That's why the weighting is log-discounted rather than linear — the same shape used in search relevance metrics, for the same reason: attention drops off steeply.

Citation

Did the engine link to your site as a source? This is different from being named, and stronger. A mention means the model knows about you. A citation means it treated your site as a checkable authority — and for the reader, it's a click.

Sentiment

How were you described? "The best option for small teams" and "cheap, but limited" are both mentions. They are not the same outcome, and a score that treats them identically is measuring exposure rather than reputation.

Putting them together

A composite works like this. If you're not mentioned, the response scores zero — there's no partial credit for absence. If you are mentioned, you earn a presence floor, then the remainder is distributed across prominence, citation and sentiment.

The result is a 0–100 number where moving from "named last, uncited, lukewarm" to "named first, cited, recommended" is a large, visible gain — because that is a large, real gain.

Two things that also belong in the aggregate

Not every engine is worth the same

A mention in the assistant most of your customers actually use is worth more than one in a niche engine. Weighting engines by real usage isn't favoritism; treating them as interchangeable is the distortion.

Old results should fade

A check from eight months ago is not a claim about your visibility today. If it counts at full weight forever, your score becomes a mixed-vintage average that describes no particular moment — and comparing two points on that chart means nothing.

We apply a 30-day half-life: results lose half their weight each month and stop counting after 90 days. The score answers "how visible are you lately", which is the only version of the question that's actionable.

The score you can't fake your way up

There's a side benefit to a composite. A binary presence score can be gamed by chasing obscure long-tail questions where you're the only plausible answer. Your percentage climbs; your business doesn't.

A composite resists that. Winning a question nobody asks still scores, but position, citation and sentiment on questions that matter dominate the number. The metric points in roughly the same direction as the outcome — which is the only real test of a metric.

What we changed

AEO Track's AEO and GEO scores are now composites of presence, prominence, citation and sentiment, engine-weighted and recency-decayed, reported with a confidence range.

We were already collecting every one of those signals. Position was computed and displayed. Sentiment was scored by a separate AI call on every check. Citations were tracked on their own page. All of it was then thrown away at the moment the score was calculated, in favor of counting yeses.

Fixing that didn't require new data. It required using what was already there.

Frequently asked questions

What should an AI visibility score measure?

More than whether your brand was mentioned. A useful score also accounts for how prominently you appear in the answer, whether the engine cited your site as a source, and how positively it described you. Those distinctions are the difference between a recommendation that sends customers and a name at the bottom of a list.

Why is a binary mentioned-or-not score misleading?

It treats being named first with a citation and a recommendation identically to being named eleventh with no link. That means it cannot reward the improvements that matter most, and it signals success as soon as you scrape into a list.

Why weight AI engines differently in a visibility score?

A mention in the assistant most of your customers actually use is worth more than one in a niche engine. Treating every engine as interchangeable distorts the score away from real-world business impact.

Why should older AI visibility checks count for less?

A check from months ago is not a claim about your visibility today. AEO Track applies a 30-day half-life so results lose half their weight each month and stop counting after 90 days, which keeps the score a statement about now rather than a mixed-vintage average.

Score what actually matters.

Presence, position, citations and sentiment — across ChatGPT, Gemini, Claude, Perplexity and Grok. Free for 14 days, no credit card.

Start Your Free Trial →