Skip to content
Learn · AEO fundamentals

What an AI visibility audit should contain

An AI visibility audit should contain a frozen prompt set with the run count printed on every figure, a named list of the engines it reached and the engines it could not, six pillars reported separately rather than summed into a score, an explicit diagnosis of whether you are absent because you are never retrieved or retrieved and not chosen, and a fix list ordered by that diagnosis rather than by pillar.

Vignette: a name being marked inside a generated answer. Illustration, not measurement: no figure appears in the loop.
In short

Key takeaways

  • The defect in a screenshot report is not that it was cheap to produce. It is that nothing in the artefact tells you whether the finding would happen again, because a single run of a probabilistic system is a sample of one.
  • Six pillars are only worth having if none of them averages into the others. They run on different clocks: crawler access changes in days, third-party consensus changes in quarters if it changes at all.
  • The sentence most audits omit is which of the two gates is binding. Never retrieved is an authority problem, retrieved and not chosen is an extractability problem, and buying the wrong one produces no movement and no explanation for it.
  • Engine coverage is a claim about what was not measured. A list of engines with no exclusions on it is an answer that has not been given.
  • A method with no way to return a finding you did not want has no way to return one you did. Ask any provider, us included, to show you a measurement that cost them something.

What the screenshot version leaves out

There is a version of this deliverable that takes about twenty minutes. Somebody opens a chat window, types a question a buyer might ask, screenshots the answer, and pastes it into a slide with a sentence of commentary underneath. Sometimes it is billed and sometimes it is given away. It is genuinely good at one thing, which is making an executive who has never thought about the problem care about it inside four seconds, and that is not nothing.

The defect is not that it was cheap to produce. It is that nothing in the artefact tells you whether the finding would happen again. Generative engines are probabilistic, so the same question asked an hour later can name a different set of companies, and there is no way to look at a screenshot and know which side of that variance you are seeing. We measure the size of it on one surface and it is not small: AI Overview presence for a given query moves run to run by roughly fifteen percentage points, which is why we report presence as a rate across repeated passes with the number of passes printed beside it. A screenshot has a run count of one and does not mention it.

Four further omissions follow from the format rather than from carelessness. One engine, so the result describes one retrieval stack rather than the surfaces your buyers actually use. One session, which may or may not have carried personalisation and location into the answer. No field, so the other companies named are visible but their frequency across the category is not. No denominator, so the absence has no scale, and absent from one answer looks exactly like absent from four in five answers.

The fifth omission is the expensive one, and it survives into reports that run to forty pages. A screenshot cannot say which problem it found. It shows an outcome and leaves the mechanism entirely open, and the mechanism is what the next invoice gets spent on. Taken as a prompt to go and measure something, a screenshot is a reasonable first move. Taken as the measurement, it is one draw from a distribution, labelled as a fact.

An answer enginenot captured
A thorough AI visibility audit typically reviews your brand mentions across major AI platforms, checks your technical setup and structured data, benchmarks you against competitors, and produces a prioritised action plan.
prompt
“what should an AI visibility audit include”
captured
not captured

illustrative, not a capture Every clause in that answer is a deliverable rather than a method. It says nothing about run counts, engine exclusions, denominators, or which of the two gates the audit is supposed to diagnose, which is the part that decides whether the action plan is aimed at the right problem. Written to show the shape of an engine answer. It has no session, no capture date and no sample size, so it is not evidence about any category, including yours.

Six pillars, and why they must not be summed

Our own sweep reads six pillars, and naming them is the easy part. Citation share is how often each named firm appears in AI answers to your category's questions. Website and technical AEO covers whether a machine can reach and render what matters. Earned and paid media covers the coverage and brand-search signals that feed authority, since paid does not put you in an answer directly but the brand search it generates is a real input. Reputation and consensus covers the third-party sources a model synthesises from, which means comparison pages, forums, reviews and industry media. Conversion readiness asks whether attention that arrives can turn into revenue, because a citation landing on a page that does not convert is a vanity result. Competitive benchmark sets the other five beside the competitors you nominate, so every finding is relative rather than abstract.

Which six is arguable. Another practitioner could propose a different cut and defend it, and we would not have much to say against a good one. What is not arguable is that six pillars are only worth having if none of them is permitted to average into the others, and that summing them into one readiness score out of a hundred destroys the single property that made six better than one.

They are not commensurable, and the clearest evidence is that they run on different clocks. Crawler access changes in days. Page structure changes in weeks. Reputation and consensus, the sources a model has already absorbed into its view of your category, change in quarters if they change at all, which is why it is the likeliest explanation for a stubborn absence and rarely the first thing to buy. A score that adds a fast cheap pillar to a slow expensive one and prints one number has encoded a prescription without stating it, because a weight is a claim about how much a given problem matters.

A competitor's own documentation makes this argument, which is worth considerably more than us making it. GrowthX publishes public product docs describing how it scores pages, and read on 2026-08-15 the headings there include Health and Quality stay separate on purpose, Multiple axes separate problems a single score merges, and Incomplete is not a bad score. That is a company with no reason at all to help our position arguing that separate axes exist precisely because a single score merges problems that are different. So when you are evaluating an audit, the pillar count is not the thing to check. Whether the pillars are reported separately is.

The sentence most audits do not contain

Citation is decided in two stages, and the audit's job is to say which stage you are stuck at. The first stage is whether you are in the pool of sources an engine is willing to draw on for a question at all, which is governed by authority: rank in the index the engine grounds on, the class of host you sit on, technical eligibility. The second is whether you are selected out of that pool on any given run, which is governed by extractability: whether a coherent passage can be pulled from the page, whether the heading matches the sub-question, whether what you say corroborates what the other eligible sources say.

The fixes have almost nothing in common, so an audit that hands over a ranked inventory of everything wrong has stopped one step short of the deliverable. What a buyer needs is a sequence, because structural work on pages an engine never fetches cannot produce a measurable result at any level of craft, and no amount of authority work will change which of six eligible sources gets named on a particular roll.

One case from our own delivery makes the distinction concrete. On an account we audited, we tested four buyer questions and an AI Overview fired on all four, with no competitor holding the ground. That reads as an open surface and an obvious content opportunity. The account already had articles addressing two of those four questions. Neither was indexed. The wider check, run page by page against Search Console on 2026-08-01, found 26 of 52 pages indexed and only 3 of 22 articles. The correct finding was therefore not that the account needed content: the content existed and was not eligible, which is a pool problem with a technical fix and a fast clock. An audit reporting the AI Overview gap without the indexation check would have sold a content programme against a crawling failure. Both reports would have contained the same true observation, and only one of them would have been worth acting on.

Coverage is a claim about what was not measured

Engine coverage is the section where a report quietly becomes a claim about what it did not look at, and it is the easiest omission to hide, because a list of engines looks the same whether or not something is missing from it.

Ours, stated in full. We reach three surfaces. ChatGPT and Gemini are rendered by DataForSEO's LLM scrapers, and Google AI Overviews come from a SERP request with the asynchronous answer block explicitly expanded. We do not scrape those consumer products ourselves; the vendor renders them and carries that relationship. Describing a bought capability as an in-house pipeline would be the first small lie in a document whose entire value is that it contains none.

Two engines we do not cover, with the reasons attached. Claude has no ground-truth path anywhere: the LLM scraper supports ChatGPT and Gemini only, and the vendor API is model-level, which we filter out of sweeps structurally so that an environment variable cannot quietly re-enable one. That coverage is lost, and we would rather name the loss than estimate around it or drop the engine from the denominator. Perplexity has no scraper either, its API surface is model-level and throttled, and automating the consumer site is disallowed in its robots.txt, as it is for the other two consumer products. Being technically able to reach a surface is not permission to reach it, and a provider quoting you a Perplexity figure should be asked which of those three paths produced it.

There are limits inside the engines we do cover, and those belong in the report too. Our Gemini path emits no search results field at all, in every run we have measured, so grounding there is derived rather than observed and we never describe it as retrieval-verified. The AI Overview instrument proves that an overview was present; it cannot prove one was absent, because an empty return can be a genuine absence, an unexpanded block or a rate limit, and the three are indistinguishable in the output. And any figure produced by calling a vendor API with web search switched on is measuring a surface no buyer uses: an API retrieves, formats and cites differently from the consumer app, so labelling that number AI visibility describes an experience nobody is having.

Every percentage arrives with its denominator

A percentage without its denominator and its field definition is not a measurement, and the fastest way to see why is to watch one capture produce four defensible numbers.

We run local visibility on a 488-point grid: eight service keywords across sixty-one towns in New Hampshire, Massachusetts and Maine. On 2026-08-09 a client held a local pack position at 10 of those 488 points, which is 2.0 percent. A pack fired at all on 412 of the 488, so measured against the points where the format was even available, the same presence reads 2.4 percent. Organic presence was far broader, 54 of 488 or 11.1 percent, and it sat at positions 22 to 30, which is presence no human being will ever see. On the second run of the same grid, on 2026-08-14, one competitor held a pack position at 38.1 percent of all 488 points, with the next largest sixteen points behind at 21.7 percent.

The client being thin in the pack is the least interesting reading of that capture. The finding is that the format is available almost everywhere in the market, that one competitor has taken more than a third of it, and that the distance between the leader and the runner-up is eight times the client's entire position. None of that survives the sentence you appear in 2 percent of local packs, which is what a report without denominators would have printed.

We got that number wrong ourselves first, and that is the more useful half of the story. Re-deriving the grid on 2026-08-15 we read the pack column as a boolean and reported zero presence. It is a rank, one, two or three, and the corrected figure is 10 of 488. A misread of a single column produced a finding that was wrong in the direction that would have justified buying more work, and the only reason it never left the building is that the re-derivation was run against the raw capture file rather than against the previous document. An audit method that reads its own prior documents instead of its own raw data will reproduce every error it has ever made, with rising confidence each time.

The instrument's own ceiling belongs in the report as well. On another account, the Search Console summary export capped at 1,330 queries while the full pull returned 5,983, so roughly four in five of that site's query appearances were invisible in the export most audits are assembled from. The same account showed 265,242 impressions, 930 clicks and an average position of 23 across ninety days, read on 2026-08-15. An audit built on the capped export would have described a materially smaller company than the one that exists, and it would have been internally consistent the whole way through.

A figure is only ever comparable to itself

Comparability is the property a buyer actually needs, because an audit is the first reading of something that will be read again in ninety days, and almost nothing about these figures is comparable by default.

Two prompt sets of the same size are not the same basis. On one account our first and second sets were both 55 prompts and shared 17 of them. A percentage computed over each would sit on the same axis of the same chart and mean two different things, and nothing in either number would say so. So the set is frozen, versioned, and stamped onto every run at the moment it executes, and each period is computed inside the set that was live when those runs happened. Our tracking refuses to draw one trend line across a cohort change, which is an unpopular behaviour and the correct one.

That failure mode is not hypothetical. Reading current archive flags back against historical runs, which is the obvious way to build a chart and the wrong one, produced a 48-point swing in one account's favour on a single day. Nobody touched a figure. The basis moved underneath it.

Which leads to the question worth putting to any provider, us included: what are you not showing me. Our answer is specific. We hold any citation figure that has not run at least three times, because a single-sample reading of a probabilistic system carries a margin wide enough to invert the finding. That rule is the reason this site publishes no client citation share at all today, and it is the reason we have no revenue-attributed win to show you either. Both are real weaknesses. We would rather state them than write around them with softer verbs.

A method that can produce a finding you did not want

The most useful thing an audit method can do is return a result the person running it was hoping not to see, and that is the property to test for, because a method with no way to produce a bad answer has no way to produce a good one either.

Ours has done it. On 2026-08-10 we changed page titles for a client and the homepage's average position fell from 16.8 to 34.5 overnight, with homepage impressions dropping from 889 a day to 538. What lets us attribute that to ourselves rather than to a Google update or to seasonality is the control cohort: 34 pages we had not touched moved 0.8 positions over the same window, which is nothing. The mechanism was legible at term level too. A head term combining the service and the city fell from position 5 to 43 for the homepage while the page dedicated to that service rose from 36 to 15. Google had reassigned the term to a page that ranked worse for it. We rolled the titles back and published the whole thing.

That entry is more useful to somebody deciding whether to hire us than any win in our files, and it is worth being exact about why. It has a control group, so the finding is separable from noise. It has a mechanism, so a competent reader can argue with it. It is dated, so it can be checked against anything else that moved that week. And it cost us something to publish, which is the part that cannot be faked. A case study missing those four properties is a claim, and the right response to a claim is to ask how it was measured, which is what this entire essay is about.

The questions to put to us

Here is the checklist, written so you can use it against us. Ask what the prompt set is, whether it is frozen, and whether you are allowed to see it. Ask how many times each prompt was run, and specifically what the run count is on the figure currently on screen. The window should be given as dates rather than as a month name. Ask which engines were reached, which were not, and why not, and treat a list carrying no exclusions as an answer that has not been given. Every percentage needs a rate-or-union answer, since a cumulative union rises with time whether or not anything improved, and it needs its denominator and its field definition, including the flattering ones. Ask which of the two gates the audit diagnosed, pool or selection, and what that implies about sequence. Ask what is being withheld and why. Then ask to see a finding that cost the provider money.

None of those nine is proprietary and none is hard to answer if the work was done. A provider who answers all of them is not thereby correct, but their numbers can be checked, and a number that can be checked belongs to a different class of object from one that cannot. Our own answers are on the method disclosure at the foot of every page here, this one included.

There is a tenth question and we would put it first. Ask what result would make them tell you not to buy anything. Sometimes the honest finding is that a category is not yet being answered by AI in a way that names businesses, and that the money belongs somewhere else this quarter. We have said that, and we would rather say it on a free call than have you discover it in month three of a retainer. A diagnostic seller who cannot describe the finding that ends the sale is not selling a diagnosis.

Questions

Questions, answered plainly.

Is a ChatGPT screenshot worthless?

No, and it is good at one thing, which is making somebody who has never thought about the problem care about it in four seconds. Its defect is not that it was cheap to produce, it is that nothing in the artefact tells you whether the finding would happen again. A generative engine asked the same question an hour later can name a different set of companies, so a screenshot is a sample of one with no run count, one engine, one session, no competitive field and no denominator. Used as a prompt to go and measure, it is a reasonable first move. Used as the measurement, it is not one.

How many engines should an audit cover?

The number matters less than whether the exclusions are named. Engine coverage is a claim about what was not measured, and a list of engines with nothing marked as not covered is an answer that has not been given. We reach three surfaces: ChatGPT and Gemini through DataForSEO's LLM scrapers, and Google AI Overviews through a SERP request with the asynchronous answer block expanded. Claude has no ground-truth path anywhere and that coverage is lost. Perplexity has no scraper, its API is model-level, and automating the consumer site is disallowed in its robots.txt. We name both gaps rather than dropping them from the denominator.

Should the audit give me a score out of a hundred?

Not as the deliverable. Six pillars are only worth having if none of them averages into the others, and summing them into one readiness score destroys the single property that made six better than one. They are not commensurable: crawler access changes in days, page structure in weeks, and third-party consensus in quarters if it changes at all. A score that adds a fast cheap pillar to a slow expensive one has encoded a prescription without stating it, since a weight is a claim about how much a problem matters. GrowthX's own public product docs argue the same point, with headings read on 2026-08-15 including Multiple axes separate problems a single score merges.

What is the one thing that separates a real audit from a sales document?

Whether it says which of the two gates is binding. Citation is decided in two stages: entering the pool of sources an engine will draw on, which is an authority problem, and being selected out of that pool on a given run, which is an extractability problem. Their fixes have almost nothing in common. On one account we tested four buyer questions, an AI Overview fired on all four, and the client already had articles for two of them, neither of which was indexed; 26 of 52 pages and 3 of 22 articles were indexed overall. The correct finding was a technical eligibility problem, not a content gap, and a report without that check would have sold a content programme against a crawling failure.

Your report is free. Do the rules relax?

The withholding rule is the test of that, and it does not move. Any citation figure that has not run at least three times is held rather than published, because a single-sample reading of a probabilistic system carries a margin wide enough to invert the finding, and that rule is why this site publishes no client citation share today and no revenue-attributed win either. The same discipline produces the outcome a paid diagnosis rarely reaches: sometimes the honest finding is that a category is not yet being answered by AI in a way that names businesses, and the money belongs somewhere else this quarter. We would rather say that on a free call than have you discover it in month three of a retainer.

Method

What is measured, and what this page is not.

This is an explainer. It carries no figures, and it is not a reading of your category. The disclosure below states the instrument that produces the numbers the essay refers to, so the distinction is on the page rather than assumed.

instrument
Caul
what was measured
Nothing on this page. Where the essay refers to citation share, that figure is produced separately, per account.
how
A prompt set written once for a category and then frozen, run against every engine in clean sessions, with each answer stored unmodified.
over what window
Reviewed on 2026-08-16. The engines change, so read the essay against the date on the byline.
what this cannot tell you
An explainer is not evidence about your category. Being named is not being recommended, and it is not traffic or revenue. Any figure about your own visibility has to come from a capture of your own category, carrying its sample size and its window.
Keep reading

The rest of the cluster.

All Learn guides · Glossary

See where you stand.

The audit is a real sweep of your category, benchmarked against competitors you name, delivered on a call so the findings get explained rather than emailed. You keep the report and the underlying data whatever you decide afterwards.

Get your free auditWhat is in the report