← Back to Blog

One answer per engine: why we stopped averaging and what the ± means

Scoring version 6 asks each engine once per question and shows a range with the score. What the three-sample experiment showed, how the ± is computed, and where the method still has limits.

Since 27 September 2026 AEOTrack asks each engine each question once per check, and the score carries a ± range. Before that we took three answers per engine and averaged them. This post is the story of why we stopped, what the ± means, and what the method still cannot tell you. The arithmetic is on the methodology page; this is the reasoning behind it.

What we did before

The idea was standard. Language models sample, so sample them. Three answers per engine per question, averaged, with a range from the spread. It looked rigorous. It cost three times as much and tripled the time each check took.

What the three samples showed

We looked at the stored answers. Asked the same question three times in the same second at temperature 0.2, ChatGPT returned three byte-identical answers: same hash, same token count. Gemini's three differed in wording and scored the same. The three samples were copies. Averaging copies changes nothing, and counting them as three observations made the range narrower than the data justified. We were paying for false precision.

The variation that matters happens over days, not seconds. The engines search the live web again next week and rank a different list. A same-second resample cannot see that. A weekly re-check can.

What changed in version 6

  • One live answer per engine per question, web search on, temperature 0.2 where the engine accepts one.
  • Questions where no engine names any business (advice-only) are left out of the headline score and counted beside it, with a Replace action.
  • A link alone is not a mention. An answer that only cites your URL, without writing your name, counts toward citation, not toward being named.
  • Lists written as prose are ranked. "The main options are A, B and C" ranks A first.
  • Cut-off and errored answers are excluded whether or not they named you. Until version 5 a cut-off answer that named you was kept and one that did not was dropped, which favoured the engines that cut off most.

Dropping to one sample left scores unchanged and widened the ranges to what the real sample size supports.

How one answer becomes a score

Not named is 0. Named is 25 plus up to 75 for how you were named: 45% prominence, 30% citation, 25% sentiment. Each question is then the engine-weighted mean of its newest answers, and the website score is a recency-weighted mean of the questions. The weights:

EngineWeightAvailability
ChatGPT1.3every plan
Gemini1.1every plan
Google AI Mode1.1Pro and above, add-on on Starter
Claude0.9every plan
Perplexity0.9every plan
Grok0.7every plan

The weights are judgement calls about where people ask, not measured audience shares. Nobody publishes reliable per-engine usage, so the order is our reading of the market and the gaps between the numbers are round. That is a limit, and it is written on the methodology page rather than hidden in code.

A worked example from that page: 5 questions on 5 engines is 25 answers. Two of them name you, as a bare name-drop with no list position, no citation and no rated tone. Each of those scores 25 + 75 × 0.35 = 51.25. The other 23 score 0. After the engine weights, the website score is 5 out of 100, with a mention rate of 8%. A first score in single digits is normal, and getting named on a third question is worth more than perfecting the two you have.

What the ± means

The ± is the range to expect between two checks. A move inside it is normal variation, not a change in your visibility. Three estimates are computed and the widest is shown:

  1. A bootstrap over your questions. The question set is resampled 2,000 times with each question's answers intact and the score recomputed. This measures how much the number depends on which questions happen to be in the set.
  2. A bootstrap within each question and engine cell. With one answer per cell this is zero wide; it is kept for accounts with older multi-sample checks.
  3. A floor from the mention rate. A Wilson interval on how often you were named, rescaled onto the score. At zero mentions the floor uses what a first mention would earn, so a 0 is never shown as "0 ±0".

The band is centred on the score, clamped to 0 to 100, and computed with a fixed seed, so the same answers always give the same band. Alerts fire only when a move is at least your threshold and the two ranges no longer overlap.

What the data says about run-to-run noise

Two numbers from the 27 to 29 September 2026 reviews, measured on the platform:

  • Run-to-run flip rate: 3.4%. Between one run and the next, that share of results flipped with nothing else changed.
  • Engine error rate: 34% before the retry fix, under 0.3% after it. Errored calls are excluded from the score, so a high error rate had been quietly shrinking sample sizes.

A 3.4% flip rate is why one check is a reading and the trend is the measurement. It is low enough that a real move on several questions is visible within a week, and high enough that a single row changing on its own means nothing yet.

When no score is shown

Fewer than three usable answers: no score, and the app says "not enough data yet". No usable answer in 30 days: the last score stays readable, marked stale, until a check runs. Every question advice-only: the count is shown and nothing is scored. None of these states is averaged into anything.

The limits we know about

  • The range describes the sample you have. It does not predict next week's answers; that is what the trend is for.
  • Sentiment is a model's judgement of tone on the passage that names you. The app says so wherever it shows it.
  • Every call is an API call with no history and no location. Personalised answers in a logged-in chat can differ from what we read.
  • Google AI Mode is measured only on plans that include it and is not in the trial.
  • The score is a 0 to 100 index, not a percentage of anything. "Named in 2 of 25 answers" is the number to quote to a colleague; the score is for comparing weeks.

Every stored score keeps the version it was computed under and the trend chart marks the step, so a method change never looks like a visibility change. The earlier post Why your AI visibility score keeps changing has the background on sampling noise, with its own correction note, and What an AI visibility score should measure explains the recency weighting.

See your own reading

The free audit asks two engines one question your buyers ask and shows who they named. Free, no account, about a minute. Then track five engines weekly with a score, its range and the "How this score works" panel in the app, on a 14-day trial, no card. Monthly counts of answers read, cited and named will be on the Named Report page from the first month.

See what the engines say about you.

Presence, position, citations and sentiment — across ChatGPT, Gemini, Claude, Perplexity and Grok. Free for 14 days, no credit card.

Start Your Free Trial →