Why Your AI Visibility Score Keeps Changing (And What to Trust)
Your score says 58 on Monday and 71 on Thursday, and you changed nothing in between. That is not a bug — it is what happens when a non-deterministic system is measured once. Here is how much noise is normal and how to tell it from a real move.
You check your AI visibility on Monday and it says 58. You check again on Thursday, having changed nothing, and it says 71. On the following Tuesday it's 49.
Nothing about your brand changed. What you're looking at is measurement noise, and understanding it is the difference between acting on signal and chasing static.
AI answers are not deterministic
Ask ChatGPT "what are the best project management tools" three times in a row and you will get three different lists. Not wildly different — but the order shifts, and brands near the bottom drop in and out entirely.
This is by design. Language models sample from a probability distribution rather than picking the single most likely next word every time. That's what makes them useful and fluent. It also means any single response is one draw from a distribution, not a fact.
So when a tool asks each engine once and reports the result as your score, it's reporting one sample and presenting it as a measurement.
How much noise are we talking about?
Enough to swamp anything you're likely to achieve in a quarter.
Take a well-set-up account: 10 tracked questions across 3 engines. That's 30 yes/no observations. If your true mention rate is 40%, basic sampling statistics put the standard error at about 8.9 points — which means a 95% range roughly 17 points wide in each direction.
Sit with that for a second. If your real visibility improves by 5 points — a genuinely good quarter's work — it is invisible inside a ±17 point band. Meanwhile an account that does nothing at all will regularly show 15-point swings in both directions, purely by chance.
Single-sample scores generate false victories and false alarms in roughly equal measure. Worse, they train you to distrust the number — which means you'll also ignore it on the day it's telling you something real.
Three things that fix it
1. Sample more than once
The cure for sampling noise is more samples. Asking each engine multiple times per check and averaging cuts the noise band substantially, because independent errors partly cancel. It costs proportionally more in API calls — which is precisely why single-sample tools exist.
2. Show the uncertainty
A score of 62 tells you less than 62 ±8 from 45 checks. The second version tells you how much to trust it, and it makes the sample size visible instead of hiding it. A score of 100 built from 3 observations and a score of 100 built from 60 are very different claims about the world.
The range has to describe the score you are looking at, which is less obvious than it sounds. Our score is a composite — presence, prominence, citation and sentiment, weighted by engine — so the honest range for it is not the range for any one of its ingredients. We get it by resampling: we take the samples that produced the score, draw new sets from them thousands of times, and read off the middle 95% of the scores those redraws produce. That answers the question you are actually asking, which is "if this check ran again, how different would the number be?"
At three samples per engine, resampling alone can occasionally report no variation at all — three identical answers are not proof of perfect stability. So we floor the range using a Wilson score interval on how often you were mentioned, carried onto the score's own scale. Wilson behaves sensibly at small sample sizes, where the simpler textbook formula happily produces intervals extending past 0% or 100%, which is a good sign it isn't the right tool for that job.
3. Only alert when it's real
This is the part most tools get wrong. If your alerting fires on a fixed drop — say five points — and your noise band is ±17, you will be notified constantly about a random walk.
The better rule: only raise an alert when the confidence intervals of the two measurements don't overlap. That's a real statement — "these two numbers are different in a way sampling noise doesn't explain" — instead of a threshold that fires on weather.
Reading a score with a range on it
This is also why we suppress the score entirely below a minimum number of observations rather than showing a confident-looking figure built on almost nothing. "Not enough data yet" is more useful than a precise-looking wrong answer.
Why a noisy score is worse than no score
A number that moves randomly doesn't just fail to inform you — it actively misinforms. You'll attribute a random 12-point rise to last week's content push and do more of it. You'll attribute a random 12-point fall to a competitor's campaign and panic.
Both conclusions are drawn from noise, and both feel like insight at the time. The confidence range is what stops that happening.
What we changed
AEO Track now takes three independent samples per engine on every check, grouped under a single run. Every score carries a confidence interval and its sample count, both shown in the interface. Alerts fire only when intervals separate. And where there isn't enough data for a meaningful number, we say so rather than inventing one.
Correction, August 2026. The first version of this shipped with the range computed from how often you were mentioned rather than from the composite score it was printed beside. The two are different quantities on different scales: a score of 57 built from mentions in 11 of 12 responses was shown as “57 ±17”, when the ±17 belonged to the 91.7% mention rate. That also made drop alerts far too quiet, because the mention-rate range barely moves while the composite does. Both are fixed: the range is now resampled from the score itself, and scores measured before the fix show no range at all rather than the wrong one.
The score is less flattering as a result. It's also the first version of it we'd be comfortable asking someone to make a decision with.
Frequently asked questions
Why does my AI visibility score change when I haven't done anything?
Language models sample from a probability distribution rather than returning one fixed answer, so asking the same question repeatedly produces different results. A tool that queries each engine once per check reports a single random draw as your score, which is why it drifts on its own.
How much random variation is normal in AI visibility scores?
With around 30 observations and a true mention rate near 40%, the 95% confidence range is roughly 17 points wide in each direction. That means swings of 15 points are common with no real change, and genuine improvements smaller than that are invisible without more sampling.
What does the plus-or-minus range on an AI visibility score mean?
It is a confidence interval showing how much the score could vary from sampling alone. A narrow range such as plus or minus 5 means the number is trustworthy. A wide range such as plus or minus 19 means there is not enough data yet to draw conclusions.
When should an AI visibility drop trigger an alert?
Only when the confidence intervals of the two measurements do not overlap. Alerting on a fixed threshold such as a five-point drop produces constant false alarms when the underlying noise band is much wider than that threshold.
Get a visibility score you can actually act on.
Three samples per engine, confidence ranges on every number, and alerts that only fire when something real changed. Free for 14 days, no credit card.
Start Your Free Trial →