Skip to content
Learn · AEO fundamentals

AI visibility agency or AI visibility tool: which one you need

An AI visibility agency is hired to measure how often AI engines name a brand and then to change it, and the choice between hiring one and buying a tracking tool turns on a single question: whether anyone on your side will act on the number, because a number is cheap to buy and acting on it is not.

Vignette: a name being marked inside a generated answer. Illustration, not measurement: no figure appears in the loop.
In short

Key takeaways

  • Tracking software in this category published subscription tiers starting at $99 a month in August 2026. If a dashboard is what you need, an agency is an expensive way to buy one.
  • A visibility percentage is a fraction, and the provider chooses the denominator. Two honest methods can report very different numbers for the same brand on the same day.
  • We shipped the denominator error ourselves. Until 2026-06-02 our leaderboard labelled a cumulative union as visibility, so one internal account read 75 percent where the corrected per-run rate was 38.2 percent.
  • AI Overview presence moves by roughly 15 percentage points between runs on the same query, so presence is a rate with a run count attached, never a yes or no from one screenshot.
  • A single composite AI visibility score fuses problems with opposite fixes. One competitor's own public product documentation argues the same thing, which is corroboration from a company with no reason to help us.

What a visibility number costs to buy

Start with the price of the thing you are actually considering. In August 2026 we audited six AEO and GEO providers, and the self-serve end of that set publishes subscription tiers of $99, $299 and from $499 a month, with execution sold on top as per-unit credits. Another publishes a from-$6,000 monthly price for a platform your own team operates with a strategist in the loop. There are also free properties in the market: one of them publishes category-level AI visibility rankings at no cost, and it runs a category covering AEO and GEO providers, which means our own market is being scored daily by a competitor's free product. All of those figures are what those providers publish about themselves, captured on 2026-08-15.

So the raw number is close to a commodity. What a dashboard sells you is a percentage, refreshed on a schedule, next to a competitor list. What it cannot sell you is the judgement about whether that percentage means anything, and it cannot sell you the work that changes it.

The rest of this essay is about the first of those, because the second is obvious. If nobody on your side is going to rewrite a page, earn a mention or fix an indexation problem in response to the chart, the subscription is a monthly payment for a feeling. That is a description of what software is rather than a criticism of it.

An answer enginenot captured
Tracking tools start at low monthly prices and report brand mentions across assistants. Agencies add strategy and execution. Provider A is a self-serve platform, Provider B pairs software with a strategist, and Provider C is a managed service.
prompt
“should I hire an AI visibility agency or just use a tracking tool”
captured
not captured

illustrative, not a capture The comparison that decides this is not tool against agency. It is whether anyone will act on the number, and whether the number is a per-run rate or a cumulative union, because those two quantities can differ by tens of points on identical data. Written to show the shape of an engine answer. It has no session, no capture date and no sample size, so it is not evidence about any category, including yours.

The number is only as good as its denominator

Every AI visibility figure is a ratio. The numerator is the part nobody controls: how often an engine named you. The denominator is chosen, and that is where the meaning lives.

Two choices produce very different numbers from identical captures. A per-run rate asks in what fraction of executed prompt runs you were named. It converges as you add runs, and it falls when you are absent, which is what a reader expects a visibility figure to do. A union asks on how many prompts you were named at least once across all runs. A union can only rise. Every extra run is another chance for a hit and no chance for a miss, so it climbs toward one hundred percent as a function of how long the tool has been switched on.

We shipped exactly that error in our own product, and it is the reason we are confident others have too. Until 2026-06-02 our leaderboard computed the union across all runs and labelled it visibility. One internal account read 75 percent on that basis where the corrected per-run rate was 38.2 percent. Nobody gamed anything; the metric simply answered a different question from the one its label claimed. We renamed the union to reach, made the per-run rate the headline, published both definitions with the sample size beside them, and the regression that proves the fix is that the rate now holds steady as runs accumulate while the old union climbed with every one.

The prompt set is the other half of the denominator, and it moves the number just as hard. Drop the prompts you never win and share rises. Add prompts containing your own brand name and it rises again. Two prompt sets of the same size are not the same basis: one client's first and second sets were both 55 prompts and shared only 17 of them, so a month-over-month line drawn across that change would have been a rewrite rather than a trend. Reading current classification flags against historical runs once produced a 48-point swing in one client's favour on a single day, which is why every figure we compute is pinned to the prompt-set cohort that was live when the runs executed.

None of this is exotic. It is the ordinary reason two vendors report different numbers for the same brand, and before comparing two figures a buyer should compare the four things underneath them: the prompt set, the run count, the engines, and whether the figure is a rate or a union. Usually one of the four explains the entire gap.

Variance, and why a screenshot is not a reading

Engines do not return the same answer twice. That is not a defect in the measurement, it is the nature of the thing being measured, and it has three consequences a buyer should insist on.

First, presence is a rate. AI Overview presence for the same query varies by roughly 15 percentage points between runs in our own measurement, so the only defensible report is a rate over repeated passes with the number of passes printed next to it. A vendor screenshot showing an overview, or showing none, is one sample of a noisy process.

Second, absence cannot be proved by one instrument. Our capture proves an overview was present. An empty return is consistent with a genuine absence, an unexpanded asynchronous block and a rate limit, and those are indistinguishable in the output. Anyone reporting you are absent from AI answers, on the strength of one tool's empty result, is reporting a property of their tool.

Third, a figure below a certain sample size should not travel at all. Our rule is that a citation figure has to have run at least three times before it leaves the building. That rule is expensive and we apply it to ourselves: every client citation share we currently hold is a single-sample reading with roughly 25 points of margin, which is why there is no client citation percentage anywhere on this site. What we can show instead is measurement with its denominator attached. For 603 Basement Solutions, who have cleared us to use their name, a 488-point local grid across 8 service keywords and 61 towns showed them holding a local pack position at 10 points, or 2.0 percent, on 2026-08-09, while one competitor held 38.1 percent of the same grid. That is a diagnosis you can check the arithmetic on.

One score is the wrong shape for the problem

Most visibility products resolve to a single headline number, sometimes a branded score out of one hundred. That shape is convenient and it destroys the information you are paying for.

Citation happens in two stages with opposite fixes. Entering the pool of sources an engine will retrieve for a question is governed by authority: rank, host class, corroboration. Winning the selection once you are in that pool is governed by structure: whether the answering sentence exists as a clean passage, whether the page is the shape the engine tends to quote. A brand sitting at fifty percent is either in the pool and losing the roll, which is a rewriting problem, or marginal to the pool, which is an authority problem. One blended number cannot tell you which, and prescribing off it is how a provider sells the wrong quarter of work with a clear conscience.

We are not the only people who think so, and the corroboration is worth more than our own argument because it comes from a company with no reason to help us. GrowthX publishes its product documentation publicly, and the page on how pages are scored carries the headings Health and Quality stay separate on purpose, Multiple axes separate problems a single score merges, and Incomplete is not a bad score. That is a direct competitor arguing, in its own docs, against the composite it would be commercially easier to sell.

The practical version for a buyer: ask any provider to show you the two numbers separately. How often are you in the retrieved set at all, and how often are you selected when you are. If the product cannot produce those separately, it is a chart rather than a diagnostic, whatever the score is called.

The routing question, and where we lose

Buy the tool if you have people who will act on it. An in-house content or search team, a marketing lead who owns the channel, and an existing habit of shipping changes on evidence. In that situation a subscription plus your own judgement beats an agency retainer on price and on speed, and one of the platform providers publishes a from-$6,000 monthly price for exactly that arrangement.

Hire a service if the measurement is going to sit unread. Most owner-led businesses we speak to are in this position: the number is not the constraint, the doing is. In that case the useful engagement is the diagnosis plus the execution, and the tracking is an input rather than the product.

There is a third answer that providers rarely give, so we will. Buy nothing yet if AI answers do not name businesses in your category. That is a measurable condition and it takes a few hours of instrument time to settle. Our free assessment exists partly to produce that conclusion, and the honest close on a diagnostic is that if the channel is not live in your market we say so and you spend nothing.

Where we lose, plainly. We have no revenue-attributed win yet. We hold every client citation share we have measured until it re-runs at three passes or more, so we cannot show you the number a dashboard would show you tomorrow. We do not measure Claude at all, because there is no ground-truth path to the consumer product, and we do not measure Perplexity either. We publish no guarantee while two of the six providers we audited do. We do not have the procurement documentation an enterprise buyer expects. And we will not post through persona accounts to manufacture the appearance of being recommended, which one provider in the audited set sells as a feature. If any of that is disqualifying, it should be, and the fastest way to find out is a call that costs nothing.

Questions

Questions, answered plainly.

What does an AI visibility agency do that a tracking tool does not?

A tool reports a percentage on a schedule. An agency is accountable for changing it, which means diagnosing whether you are missing from the retrieved pool or losing the selection inside it, and then doing the corresponding work. If you have an in-house team that will act on the data, the tool plus your own judgement is cheaper and faster. If the dashboard will go unread, the subscription buys a feeling.

Why do two AI visibility tools report different numbers for the same brand?

Because they compute different fractions. Different prompt sets, run counts and engines, and in some products a cumulative union across runs rather than a per-run rate. A union can only rise, because every extra run is another chance for a hit and no chance for a miss. We made that exact error ourselves: one internal account read 75 percent on the union where the corrected per-run rate was 38.2 percent, and we renamed the union to reach on 2026-06-02.

How many runs does an AI visibility measurement need?

More than one, and our own floor is three before a figure travels. AI Overview presence on the same query moves by roughly 15 percentage points between runs, so a single capture is one sample of a noisy process. Every client citation share we currently hold is a single-sample reading with roughly 25 points of margin, which is why none of them appear on this site.

Can a tool prove my brand is absent from AI answers?

No. Our instruments prove an overview or a mention was present. An empty return is consistent with a genuine absence, an unexpanded asynchronous block and a rate limit, and those look identical in the output. A report of absence based on one tool's empty result is a statement about the tool.

Is a single AI visibility score out of 100 useful?

It is the wrong shape for the problem. Entering the pool of sources an engine retrieves is an authority problem; being selected once you are in the pool is a structural one, and they have opposite fixes. One competitor's own public product documentation makes the same argument, with headings stating that separate axes exist because a single score merges problems that are different. Ask any provider to show the two numbers separately.

Method

What is measured, and what this page is not.

This is an explainer. It carries no figures, and it is not a reading of your category. The disclosure below states the instrument that produces the numbers the essay refers to, so the distinction is on the page rather than assumed.

instrument
Caul
what was measured
Nothing on this page. Where the essay refers to citation share, that figure is produced separately, per account.
how
A prompt set written once for a category and then frozen, run against every engine in clean sessions, with each answer stored unmodified.
over what window
Reviewed on 2026-08-16. The engines change, so read the essay against the date on the byline.
what this cannot tell you
An explainer is not evidence about your category. Being named is not being recommended, and it is not traffic or revenue. Any figure about your own visibility has to come from a capture of your own category, carrying its sample size and its window.
Keep reading

The rest of the cluster.

All Learn guides · Glossary

See where you stand.

The audit is a real sweep of your category, benchmarked against competitors you name, delivered on a call so the findings get explained rather than emailed. You keep the report and the underlying data whatever you decide afterwards.

Get your free auditWhat is in the report