Measurement standards
Two AI visibility tools can give you different numbers for the same brand on the same day, and both can be right, because they measured different things. This page says exactly what we measure, how well, and what our numbers are not yet good enough for.
Where we stand today
In August 2026 the IAB published Measuring Visibility in the AI Era, a framework for judging whether AI visibility data is fit for the decision you want to make. It sets out seven criteria and two quality tiers: directional and decision-grade.
Measured against those criteria, what we can claim depends on the plan, because two of the seven criteria are bought rather than engineered. Four are the same whatever you pay: cadence, platform reporting, prompt type coverage and methodology disclosure all sit at decision-grade on every tier. The two that move are sample size (how many times each question is asked) and query volume (how many questions), and reproducibility follows them, because measuring run-to-run variation requires asking more than once.
The framework grades a programme at its weakest criterion, so the overall standing is Starter: below directional. Pro: directional. Scale: directional. Enterprise: decision-grade. Enterprise is the only configuration that reaches the top rung on all seven.
Scale is the clearest illustration of grading at the weakest link. It reaches decision-grade on six of the seven criteria, including query volume, and is held at directional by sample size alone: it asks each question twice where decision-grade wants three. More questions widen what you can see; only more repeats narrow the range on what you have already seen.
Inside the product a programme below the directional bar is labelled Exploratory. That word is ours, not the framework’s. IAB defines two tiers and we needed a name for the rung underneath, so wherever you see it, it means exactly what this page calls below directional.
We publish that rather than round it up, and we show the same tier inside the product next to your score.
Our standing, criterion by criterion
| Criterion | Starter | Pro | Scale | Enterprise | Why |
|---|---|---|---|---|---|
| Sample size | Below directional | Directional | Directional | Decision-grade | Starter asks each question once per engine, which is a sample of one and cannot show whether a reading is typical. Pro asks twice, the framework’s directional bar. Enterprise asks three times, which is the decision-grade bar. Wherever a question is asked more than once, each answered cell reports how often that engine actually cited you, 2 of 3 rather than a single yes or no, both in the grid and in the export. |
| Query volume | Below directional | Directional | Decision-grade | Decision-grade | 25 questions on Starter. The framework "treats fewer than 50 queries per measurement program as exploratory rather than directional", so Pro tracks 50 and Enterprise 100, which is the decision-grade floor. |
| Reproducibility | Directional | Decision-grade | Decision-grade | Decision-grade | The directional bar asks a provider to document how much variation is typical, and the ranges below do that for every tier. Decision-grade additionally requires the variation measured within the window, which means genuinely re-running your queries. Pro and Enterprise do; Starter, asking once, cannot. |
| Prompt type coverage | Decision-grade | Decision-grade | Decision-grade | Decision-grade | All four of the framework’s intent types in every question set, with results segmented by intent in the product. This used to depend on your business type, and no longer does. |
| Methodology disclosure | Decision-grade | Decision-grade | Decision-grade | Decision-grade | This page, plus the framework's disclosure fields filled in (platform coverage and model versions, prompt library, query sourcing, collection architecture, panel validity, baseline resets, and what each falls short on), per-scan provenance in the product, a full data export, and a changelog of every change to the scoring method. |
| Testing cadence | Decision-grade | Decision-grade | Decision-grade | Decision-grade | Every tracked brand is scanned on a fixed weekly schedule. The framework grades weekly or better as decision-grade, which is why we withdrew the daily option rather than charging for it. |
| Platform reporting | Decision-grade | Decision-grade | Decision-grade | Decision-grade | Six engines, reported separately and never blended into one figure: ChatGPT, Gemini, Claude, Perplexity, Grok and DeepSeek. |
How much do the answers actually move?
AI answers are not stable. Ask the same question twice and you may get a different result, so any single reading carries uncertainty. The framework asks providers to document how much variation is typical, so here is ours.
We asked the same eight questions ten times each across all six engines, in one window, and counted how often the verdict changed between consecutive runs. On a heavily cited brand the average flip rate was 4.2%, ranging from 0% on Claude to 8.3% on DeepSeek. Across our stored history, over roughly 525 comparable pairs per engine, the equivalent figure is 3.2%.
Three honest caveats. Variation depends heavily on the brand: we repeated the exercise on a barely cited domain and five of the six engines never mentioned it at all, which produces a flip rate of zero that means "invisible" rather than "stable". These ranges are also what we measured elsewhere rather than a confidence interval on your own number: Starter asks each question once per scan, so the ranges are the only guide to variation it can give you, while Pro asks twice and Enterprise three times, which measures the movement on your own questions inside the window. And both runs were made on the engines' own memory rather than with web grounding, while every paid tier runs grounded, so the grounded figures may differ.
What asking ten times actually buys
A flip rate tells you how often two runs next to each other disagreed. It does not answer the question you actually have, which is: if I had only asked once, how likely is it that the answer I was shown was not the typical one? That is a different quantity, and the same two runs answer it.
Take one question on one engine, asked ten times. Whichever verdict came back more often is the majority verdict. Now pick one of those ten answers at random, which is what a single-sample tool shows you: how often does it point the other way? On the heavily cited brand, 17 of 480 readings, 3.5%, or roughly one reading in 28. Counted by square rather than by reading, seven of the 48 question-and-engine squares moved at all, 14.6%, so about one square in seven of a one-shot grid is a square that a repeat could have turned over.
| Engine | Mention rate | Squares that moved | Single reading disagrees |
|---|---|---|---|
| Claude | 100% | 0 of 8 | 0% |
| ChatGPT | 88.8% | 1 of 8 | 1.3% |
| Gemini | 78.8% | 1 of 8 | 3.8% |
| Perplexity | 16.3% | 1 of 8 | 3.8% |
| Grok | 96.3% | 1 of 8 | 3.8% |
| DeepSeek | 91.3% | 3 of 8 | 8.8% |
travelodge.co.uk, eight questions asked ten times each on six engines in one window, 10 August 2026. Not grounded: see the caveats below.
The spread between engines is the part worth keeping. DeepSeek disagreed with itself on 8.8% of readings while Claude, which named the brand in 100% of answers, never moved at all. A tool that samples every engine the same number of times is spending the same money on a stable engine and an unstable one.
And the same figure on the barely cited brand is 0.2%, which is not good news. Five of the six engines never mentioned that brand once in 480 readings. Nothing can disagree about a brand nobody names, so invisibility scores as perfect reproducibility. Any stability number, ours included, has to be read next to the mention rate it was measured on, or it rewards being ignored.
Three things to hold against that figure. Both runs were made on the engines' own memory rather than with web grounding, and every paid tier runs grounded. Two domains, eight questions each, is a property of these two runs and not a confidence interval on your own number. And the majority of ten is itself an estimate rather than the truth; it is a far better estimate than one reading, which is the whole point, but it is not ground truth.
One disclosure about how it was produced. The two runs were recorded as per-engine summaries and the individual answers were not kept, so the per-square counts here are recovered from those summaries rather than re-read from raw data. Each published rate is a whole number of observations over a known total, so it inverts to exactly one count, and the per-square split is then forced by arithmetic. The script enumerates every arrangement consistent with the mention rate, the moved-square share and the flip rate together, and reports a range rather than a figure if more than one survives. On these two runs every engine resolves to a single value. It is tools/single-sample-disagreement.js, it calls no model and costs nothing to run, and a test re-runs it against the committed runs on every build.
The first grounded observation, and why it is not a study
That last caveat is the real gap in what we publish, so here is the first thing we have measured on the other side of it. On 11 August 2026 we ran three questions three times each across all six engines, grounded, on our own domain: 18 question-and-engine cells, all of which answered. One of them disagreed with itself across the three runs, a flip rate of 5.6%.
That is one data point, not a result. 18 cells is far too small a sample to put a range on, it is a single domain on a single day, and one cell landing the other way would move the figure by 5.6 points on its own. Nobody should plan against it, including us. It was published because, at the time, it was the only grounded number of any size we had.
The funded grounded study, which went against us
That gap is now closed. We repeated both memory-layer runs at full size with web grounding on: 8 questions, 10 repeats, all six engines, 480 calls per domain. Same questions, same repeats and same engines as their memory-layer twins, so the comparison is like for like.
Grounded measurement is about four times noisier than the memory layer. On travelodge.co.uk, a brand the engines cite readily, the mean flip rate went from 4.2% on memory to 16.2% grounded, a factor of 3.9. On searchscore.io the same comparison gives 0.5% against 2.1%, a factor of 4.5. Two domains at opposite ends of visibility, agreeing on roughly the same multiple.
The sharpest single figure: on travelodge.co.uk, 100% of ChatGPT cells disagreed with themselves, at a flip rate of 37.5%. Every question we asked ChatGPT about that brand, grounded, returned a different answer across repeats. One reading from that engine is close to a coin toss, and no amount of presentation makes it otherwise.
Most of that increase is arithmetic rather than unreliability, and saying so is not a walk-back. A yes-or-no measurement cannot disagree with itself at all when the answer is always yes, and disagrees most often when the answer is a coin flip: two independent readings at rate p differ with probability 2p(1 − p). Grounding moved this brand off the ceiling. On memory, Claude named it in 100% of readings, Grok 96%, DeepSeek 91%; grounded, those became 55%, 86% and 58%, because a grounded engine reads current pages and frequently names a competitor instead. Rates in the middle flip more whoever is measuring them.
Measured against that chance baseline, repeatability barely moved: observed flips run at 0.34 of the independent-draw rate on memory and 0.36 grounded. The engines repeat themselves about as much in both layers. One exception matters, and it is the biggest engine: ChatGPT went from 0.14 to 0.87, so its grounded repeats are close to genuinely independent draws rather than a recalled answer. That is a real change in the measurement, not just a change in the thing measured.
The practical consequence survives the decomposition intact. Whatever the cause, one grounded reading of a mid-range brand is unreliable, and mid-range is exactly where a brand worth measuring sits. An engine that names you every single time is not giving you a measurement either; it is giving you a ceiling.
This is worse for us than the figure it replaces, and that is why it is here. The paragraph this section used to carry promised to publish the grounded study whichever way it landed, and it landed badly: the memory-layer figures above were understating the noise a paying customer actually experiences, because every paid tier runs grounded. The honest reading of the four-times factor is that our own published variance was flattering, and the correction belongs on the same page as the original claim rather than in a changelog nobody opens.
It has a consequence we are not hiding either. Pinning a rate to a standard error of fifteen points needs between 6 and 12 samples per engine on the cited brand, measured grounded. No tier we sell runs that many: Enterprise runs three, Pro two, Starter one. Repeat sampling narrows the band and is the difference between a directional and a decision-grade programme, but it does not abolish the noise, and a vendor telling you a single grounded reading is a measurement is selling you a coin toss with a chart around it.
Two limits on the study itself. It is two domains, so it is a property of these runs rather than a confidence interval on your own number, exactly as the memory-layer caveat above says. And a flip rate on an engine that never mentions the brand means invisible rather than stable, which is why the searchscore.io figures sit lower: there is less to disagree about. The two runs cost $11.4881 and $12.6378, are committed to the repository as the fixtures behind every figure in this section, and are regenerated by tools/variance-study.js --grounded.
Definitions that differ between vendors
A mention is not a citation
We report both, separately. A mention is your brand appearing anywhere in an answer. A citation is the answer relying on you as a source, either by linking to you or by naming you as the source. Tools that report a single "visibility" number usually mean the first and imply the second.
The distinction is worth an argument only if the two numbers actually differ, so here is ours. Across every observation we have stored, 6,457 question-and-engine readings over 15 domains and 109 scans between 2 June 2026 and 11 August 2026, the mention rate is 24.3% and the strict citation rate is 4.4%. They differ by 5.5 times. Neither is the real number. They are two different events, and a report that gives you one of them under both words is off by that multiple in the flattering direction.
| Engine | Mention rate | Citation rate | Returned a source list |
|---|---|---|---|
| Perplexity | 30.3% | 16.7% | 65.8% |
| DeepSeek | 27.4% | 3.6% | 14.1% |
| Claude | 14.6% | 2.4% | 13.3% |
| Grok | 20.7% | 2.1% | 13.5% |
| ChatGPT | 26% | 1.7% | 11.7% |
| Gemini | 26.6% | 0% | 11.4% |
6,457 observations. 131 failed calls are excluded from both rates rather than counted as non-mentions.
Read the third column before the second. A per-engine citation rate is partly a property of the engine, not of the brand. Perplexity handed back a source list on 65.8% of its answers in this corpus and Gemini on 11.4%, a spread of 5.8 times. An engine that rarely returns a source list cannot produce a linked citation however well a brand is doing, so a low citation rate on that engine is a statement about what the engine returned, not about how the brand performed there. We will not present it as the latter and neither should anyone else.
Two further limits on that table. 98% of these readings predate the field that records whether the answer was fetched with web grounding or from the engine's own memory, and link availability depends heavily on which, so those columns describe this corpus rather than a permanent property of the engines. And our detector for a brand named as a source without a link is deliberately narrow: it accounts for only 6.6% of the citations counted, so in practice the strict rate is close to a link rate, and we would rather label it that way than let a narrow detector flatter it.
Anyone can recompute this. It is tools/mention-vs-citation.js, it reads the database read-only, it calls no model, and the classifier it uses is app/lib/citation-kind.js, which implements the IAB definition rather than one of ours.
How accurate is the mention detector itself?
Every rate above depends on one classifier deciding whether an answer mentioned a brand, so its own accuracy is a fair question. We measured it blind and out of sample on 19 August 2026: 47 stored answer cells, labelled by an annotator who has never edited the classifier, with the classifier's verdict hidden until after each choice. The labels are committed to the repository, so the measurement re-runs against any future version.
| Result | 95% interval (Wilson) | |
|---|---|---|
| Precision | 7 of 7 correct, no false positives | 64.6% to 100% |
| Recall | 7 of 8 (87.5%) | 52.9% to 97.8% |
| Agreement with the human label | 46 of 47 (97.9%) | 88.9% to 99.6% |
We report the precision row as a count, not a percentage, on purpose: 7 positives cannot carry a point estimate, and the interval is wide. Four further limits, stated rather than buried. The sample is stratified, not random: it over-weights the cells where the classifier and a naive matcher disagree, which makes it a harder test than a random draw but means these are not population estimates. There is one annotator, so inter-rater agreement is unmeasured. The single disagreement is a false negative we had already documented, independently reproduced by a second blind annotator. And the intervals overlap the pre-fix study's (66.7% precision on 126 cells), so this is consistent with the fixes having worked and is not proof of it.
What the fixes are worth is measurable a different way. Re-scoring all 5,977 stored cells, the tightened classifier removes 11 answers that previously counted as citations and adds none, moving the stored rate from 21.78% to 21.6%. Every one of the 11 was read by hand; each is an engine saying it had never heard of the brand.
Anyone can re-run the measurement: node tools/brand-mention-regress.js --fixture tests/fixtures/accuracy/brand-mention-sample-2026-08-18.json --verbose.
Questions that name you are scored separately
Asking an AI assistant about a brand by name will usually surface that brand. Counting those in a headline score flatters it, so questions that name you feed a separate brand defence figure instead of your main score. Where a buying question names you and a competitor together, we keep it in the score, because it is a real buying question, but we show you how much of your result rests on it.
A failed call is not a zero
If an engine is unreachable we record no observation. We never count an outage as the engine choosing not to mention you, because that would show as a drop you did not cause. If every engine fails, the scan is discarded rather than saved as a collapse to zero.
What each tier still does not give you
- Starter is below the directional bar on sample size and query volume, and always will be: 25 questions asked once is a spot check, priced as a spot check. Use it to see whether you appear at all, not to report a number to anyone else.
- Pro clears directional on all seven. It does not reach decision-grade, because that needs 100 questions asked three times, and at Pro's price that configuration would cost more to run than the subscription collects.
- Enterprise reaches decision-grade on all seven. There is no higher rung in the framework, and there is nothing we are holding back from it.
Two of these criteria are honestly a pricing question rather than an engineering one. Asking every question three times across six engines costs three times as much as asking once, every week, per customer, and web grounding is paid on every call rather than cached between samples. We would rather price each tier at what it can actually deliver and say so here than describe a spot check as decision-grade. The framework's own warning is that the common failure is "treating directional data as decision-grade without recognizing the gap".
Check it yourself
Every Tracker account can export the raw observations behind its score as CSV or JSON: one row per question, engine and scan, with the retrieval mode, the model that actually answered, and the measurement baseline each row belongs to. Where the question was asked more than once, the row also carries how many runs it stands on and how many of them cited you, so the frequency behind the grade can be recomputed rather than taken on trust. If you cannot recompute a number, you are taking it on trust, and the whole argument of this page is that you should not have to.
Two rows are only comparable if they were scored the same way, so each row also carries the scorer version that produced it, left blank on scans older than the versioning itself rather than guessed at. The scorer changelog lists every change we have made to the scoring method, the dates each version was in force and why it changed, and it is generated from the scoring code rather than written beside it. Inside the product, a movement measured across one of those changes is not reported as a movement.
For how the numbers are produced rather than how they grade, the IAB disclosure fields are filled in on their own page: the engine panel with the model version behind each name, how question sets are built and where they come from, how responses are collected, what the panel does and does not represent, and when a baseline resets. Each field carries what we fall short on in the same row.
A standard is only worth having if you can hold something to it. The free check will score your own site and asks you for no email to do it. It reads the site rather than the engines, so it is the step before the measurement described on this page rather than that measurement itself.
Reviewed August 2026. Framework reference: IAB, Measuring Visibility in the AI Era, August 2026.