Skip to content
Learn · AEO fundamentals

Why a single composite AEO score is a defect

A composite AI visibility score fuses signals whose fixes point in opposite directions, so the number cannot tell you which work to buy, and prescribing off it is how a quarter of budget goes to the wrong problem.

Vignette: a name being marked inside a generated answer. Illustration, not measurement: no figure appears in the loop.
In short

Key takeaways

  • The two failures behind a low score are never being retrieved, and being retrieved but not chosen. The first is an authority problem, the second is an extractability problem.
  • Their fixes are opposite. Rankings, links and entity work move the first. Structure, headings, format and passage density move the second.
  • A composite hides which of the two you have, and the weights that produced it are almost never published.
  • Separating them is cheap: hold one prompt constant, run it repeatedly, and read the size of the source pool separately from your rate inside it.

What a composite score is made of

A composite takes several measurements on different scales, applies weights to them and prints one number, usually out of a hundred. A crawl check, a schema check, a mention count and a sentiment read go in; a single figure comes out and moves month to month.

The appeal is obvious. One number is reportable to somebody who was not in the meeting. The defect is equally simple and much less discussed: the inputs cannot be recovered from the output. Two companies can hold the same score for opposite reasons, and the score gives no way to tell which is which. Worse, the weights encode a prescription that nobody stated out loud, because a weight is a claim about how much a given problem matters.

Two gates, in sequence, with opposite fixes

Citation is decided in two stages, and the stages are governed by different things. The first stage is entry to the pool of sources an engine is willing to draw on for a question. That is gated by authority in the index the engine grounds on: whether the page is indexed and eligible to be shown with a snippet at all, how it ranks, what class of host it sits on. Google is explicit that its own generative answers rely on its core Search ranking systems to retrieve pages from the Search index, which makes rank in that index the entry ticket rather than a legacy metric.

The second stage is selection from the pool. Being in the pool is not being cited. Which of the eligible sources ends up in the answer is governed by extractability: whether a coherent passage can be pulled from your page, whether the heading matches the sub-question, whether the format matches the shape of the answer, whether what you say corroborates what the other eligible sources say.

The fixes have almost nothing in common. A pool problem is fixed with rankings, links, entity work and technical eligibility, on a clock measured in months. A selection problem is fixed by rewriting structure on pages that are already being retrieved, on a clock measured in days. Buy the wrong one and you get no movement and no explanation for it.

Why the blend is expensive rather than merely imprecise

Consider two companies that both score badly. The first is retrieved most times its category is asked about and named some of the time. Its pages are in the pool and losing the selection. It buys a content rewrite and starts appearing in weeks.

The second is barely retrieved at all. It buys the same content rewrite, because the same score prescribed it, and rewrites pages an engine was never fetching. Nothing changes for two quarters, because the work addressed a stage the company had not reached. The money was spent, the work was competent, and it was aimed at the wrong gate.

That is the argument for separating the two, and it is a money argument rather than a methodological preference. A blended score sold with a prescription attached is how the expensive wrong quarter gets bought, and the buyer cannot detect the error from anything on the report.

How to tell which one you have

The diagnostic is inexpensive. Hold a single prompt constant and run it many times; thirty runs is the floor we use. Then read two things separately. Count the distinct sources cited across all of those runs, which approximates the size of the pool the engine is drawing from. Then count your own appearance rate inside that pool.

The same appearance rate means opposite things depending on the pool. A company appearing in half the runs inside a stable pool of six sources is eligible every time and losing a stochastic selection, which is a structure problem with a fast fix. A company appearing in half the runs inside a churning pool of forty sources is being retrieved intermittently, which is an authority problem with a slow one. One number, two diagnoses, and the pool size is what distinguishes them.

Nothing about that method is proprietary. It is written here so that you can ask any provider to run it, including us, and so that a report which cannot say which of the two it diagnosed can be recognised for what it is.

Where the weights come from

Ask for the weights and the derivation. In our experience of reading these products the weights are chosen rather than fitted, which makes the score an opinion carrying a decimal point. That is not automatically wrong, but it has to be declared, because a weighted average of signals with opposite prescriptions is a number whose movement cannot be interpreted.

This applies to us. Our own audit produces a score out of a hundred, and the honest statement is that the rubric is ours, we say so, and the score is a summary of findings rather than the diagnosis. The findings are the deliverable. Anyone, ourselves included, who hands you a composite and a recommendation in the same breath owes you the derivation, and if it is borrowed from a third-party checklist then the checklist's author is no more authoritative than the person quoting it.

What to report instead

Two numbers rather than one, each with its run count: how often you are retrieved into the pool at all, and how often you are selected once you are in it. Then the findings, ordered by what they cost to fix against what they would change, with the stage each one belongs to named on it.

And one sentence that most reports omit: which of the two problems this engagement diagnosed. If the answer is both, the sequence matters, because selection work on pages that are not being retrieved is work that cannot show a result. If the answer is neither, because the category does not yet produce answers that name businesses, the correct recommendation is to spend the money somewhere else, and we would rather say that on a free call than discover it in month three of a retainer.

Questions

Questions, answered plainly.

Is your own audit score not the same thing?

It is a composite, and we hold it to the rules above: the rubric is ours and we name it as ours, the score summarises findings rather than replacing them, and no recommendation is made from the score alone. Where our score and a specific finding disagree, the finding wins, because the finding is the thing that was measured.

Why does almost every product publish one score?

Because one number is sellable, comparable across accounts and easy to put on a dashboard that moves. Those are real product virtues. They are just not diagnostic virtues, and the moment a prescription is attached to the number, the missing decomposition starts costing the buyer money.

Can a company have both problems at once?

Frequently, and the order is not optional. Pool entry comes first, because structural work on pages an engine does not retrieve cannot produce a measurable result. The useful output is not a ranked list of everything wrong, it is which gate is currently binding.

What if a provider will not share the weights?

Then treat the score as a proprietary index rather than a measurement, and ask instead for the two underlying counts. A provider who can produce pool membership and selection rate separately does not need you to trust the composite, and one who cannot is asking you to.

Method

What is measured, and what this page is not.

This is an explainer. It carries no figures, and it is not a reading of your category. The disclosure below states the instrument that produces the numbers the essay refers to, so the distinction is on the page rather than assumed.

instrument
Caul
what was measured
Nothing on this page. Where the essay refers to citation share, that figure is produced separately, per account.
how
A prompt set written once for a category and then frozen, run against every engine in clean sessions, with each answer stored unmodified.
over what window
Reviewed on 2026-08-15. The engines change, so read the essay against the date on the byline.
what this cannot tell you
An explainer is not evidence about your category. Being named is not being recommended, and it is not traffic or revenue. Any figure about your own visibility has to come from a capture of your own category, carrying its sample size and its window.
Keep reading

The rest of the cluster.

All Learn guides · Glossary

See where you stand.

The audit is a real sweep of your category, benchmarked against competitors you name, delivered on a call so the findings get explained rather than emailed. You keep the report and the underlying data whatever you decide afterwards.

Get your free auditWhat is in the report