Why an AI visibility score needs a measurement window
An AI visibility figure is comparable to another figure only when both were produced by the same prompt set over a stated window, because changing the set or the window changes the number even when nothing changed in the answers.
Key takeaways
- The prompt set is part of the metric. Two figures produced by different sets are not two readings of one trend, whatever the chart does with them.
- A metric that filters on today's prompt list while scanning last month's runs rewrites months you have already reported.
- Two sets of the same size are not the same basis. Compare membership, never counts.
- The window is a choice you made. The age of the thing you are measuring is a fact about it, and a young cohort gets no verdict at all.
What does a measurement window actually fix?
Engines re-rank continuously, and they sample their own output, so two captures of the same question differ for reasons that have nothing to do with your site. A window is what turns a sequence of photographs into a comparison. It fixes which captures belong to a figure and it makes the next figure answerable to the same rules.
Two things have to be fixed, not one. The period is the obvious one: a figure covers dates, and the dates get published with it. The prompt set is the one that gets forgotten, and it is the one that does the damage, because it sits in the denominator of every ratio computed over it.
Why the prompt set is part of the number
A prompt set is a hypothesis about what your buyers ask. It is written once for a category, and the moment it is written it becomes the measuring instrument rather than a list of ideas. Every share, rate and trend computed afterwards is computed over that set, so the set is not an input to the metric. It is part of the metric's definition.
This has an unpleasant consequence for anyone who prunes. It is natural to look at a set after the first capture, see prompts that never return you, and cut them as noise. The share rises immediately, because the same numerator is now divided by a smaller denominator. Nothing improved. In one account we have run, pruning would have moved a share from the low forties to the seventies in two passes without a single new mention appearing in a single answer.
So a set gets sealed. Membership is frozen, the set carries a version, every run is stamped with the version it executed under, and a report states which version produced its figures. Pruning is allowed, and after the first baseline it is usually correct, because prompts that fire nothing for anybody are paying for runs. What is not allowed is pruning and then drawing one line across the change.
How a set change rewrites months you already reported
Our own tool did the worse version of this to us, and it is the clearest example we have of why the discipline matters. Until 2026-08-09, every metric filtered on the prompt's current status while scanning historical runs. That reads today's list against last month's data, which means archiving a prompt today changed what a past month had measured.
It fired on a live account on the day thirty-eight prompts were archived. A month that had measured 40.0 percent on the set that was live at the time read 88.2 percent afterwards. Nobody touched a metric, nobody edited a run, and the number moved forty-eight points in the client's favour on its own. If it had gone the other way we would probably have caught it sooner, which is the part worth sitting with.
The fix is structural rather than procedural. Set membership is frozen in its own table, each run records the set it belonged to, and every metric resolves against that stored basis rather than against the live list. A trend series refuses to draw a single line across a basis change instead of quietly averaging one. A rule that depends on somebody remembering is not a rule.
Two sets of the same size are not the same basis
The most convincing wrong comparison is between two sets of equal size, because equal size looks like stability. In one account we run, two versions of the prompt set both contained fifty-five prompts and shared seventeen of them. Read by count, nothing had changed. Read by membership, two thirds of the instrument had been replaced.
So the comparison is always on membership. Before any month-over-month reading, the two sets are diffed and the overlap is stated. If the overlap is partial, the honest options are to report on the intersection and say so, or to declare a new baseline and start the series again. Averaging across the change produces a line that looks like performance and is arithmetic.
The window is your choice, the subject's age is a fact
There is a second window error, and it is the one that produces confident findings out of nothing. A reporting tool offers a maximum window, the analyst selects it because more data seems better, and the resulting sentence describes the tool's window as though it were the lifetime of the thing being measured.
We did exactly that on 2026-08-14. We reported that a cohort of pages had produced no impressions in sixteen months and recommended removing them from the index. Every page in the cohort had been published inside the previous five weeks, with a median age of twenty-four days. Sixteen months was the widest window the reporting tool offers, not the life of the pages. Three to five weeks of silence is the expected state for a new section, so there was no finding at all, and the supporting evidence was invalid for the same reason.
The rule that came out of it is short. Pull publish dates before choosing a window, then report the age next to the metric. A young cohort gets a re-measure date rather than a verdict. And a uniform zero across a whole cohort is more often a broken instrument or a wrong window than a real result, which is the subject of the essay on proving absence.
What the versioning discipline looks like in practice
Written down, it is four habits. Seal the set before the first capture and give it a version. Stamp every run with the version it executed under, so a figure can be recomputed from stored runs years later. Keep a changelog of what entered and left the set and when. State the version and the dates on every report that carries a figure.
None of that is expensive, and it is the difference between a number that survives being repeated to a partner or an acquirer and one that dissolves the first time somebody asks how it was produced. It is also the only honest way to answer the question a client asks in month three, which is whether the change on the chart is the work or the instrument.
Questions, answered plainly.
How often should the prompt set change?
Rarely, and never quietly. After a first baseline it is usually right to retire prompts that returned no signal for anyone, because they are spending runs. Anything beyond that is a new baseline: version the set, keep the changelog, and start the series again rather than drawing one line across the change.
Can I compare my score to a competitor's score from another tool?
Not directly. The two numbers are ratios over different prompt sets, different run counts and different engine coverage, so the gap between them is mostly instrument. What does compare is your standing against named competitors inside one capture, because that comparison holds the instrument constant for everyone in it.
What window do you use?
We publish the window with the figure rather than fixing one number here, because the right window depends on how often the set is run and how volatile the category is. What is fixed is the rule: dates on every figure, the run count beside it, and no figure at all where the runs have not separated a reading from variance.
Why does a trend line ever refuse to draw?
Because the two ends of it were measured by different instruments. When the prompt set changed between two points, our series stops rather than interpolating across the change. A gap in a chart is honest; a continuous line across a basis change is a claim nobody measured.
What is measured, and what this page is not.
This is an explainer. It carries no figures, and it is not a reading of your category. The disclosure below states the instrument that produces the numbers the essay refers to, so the distinction is on the page rather than assumed.
- instrument
- Caul
- what was measured
- Nothing on this page. Where the essay refers to citation share, that figure is produced separately, per account.
- how
- A prompt set written once for a category and then frozen, run against every engine in clean sessions, with each answer stored unmodified.
- over what window
- Reviewed on 2026-08-15. The engines change, so read the essay against the date on the byline.
- what this cannot tell you
- An explainer is not evidence about your category. Being named is not being recommended, and it is not traffic or revenue. Any figure about your own visibility has to come from a capture of your own category, carrying its sample size and its window.
The rest of the cluster.
See where you stand.
The audit is a real sweep of your category, benchmarked against competitors you name, delivered on a call so the findings get explained rather than emailed. You keep the report and the underlying data whatever you decide afterwards.