How to evaluate an answer engine optimization agency
Answer engine optimization agency, AEO agency, GEO agency and LLM SEO agency name the same service, and because the deliverable is a number rather than a ranking you can look up, the only reliable way to compare two providers is to audit the method behind their figures: the prompt set, the run count, the window, the engines covered, and what they decline to publish.
Key takeaways
- The four category names are a keyword split, not four services. One provider we inventoried runs three separate URLs for one AI-search offering, one per search term.
- Audit the arithmetic before the agency. A published ranking whose values are impossible under the provider's own stated method tells you more than any case study.
- Ask what the provider withholds. A method that never produces a held figure is not a method, it is a chart.
- Ask for the control group. Without an untouched cohort, a movement after a change is a coincidence with a date on it.
- The five things a credible engagement contains: a measured baseline, named competitor benchmarking, shipped execution rather than a dashboard, re-measurement against that baseline, and the methodology shown.
Four names, one service
Answer engine optimization is the formal version of the term. Generative engine optimization, AI search optimization and LLM SEO are the same offer in different vocabulary, and the market knows it. When we inventoried Breaking B2B's 222 URLs from their own sitemap in August 2026, we found three separate service pages for a single AI-search offering: one at an answer-engine-optimization URL, one at a generative-engine-optimization URL and one at an LLM-SEO URL. That is a keyword split, done deliberately because buyers arrive typing three different phrases. It is cheap, it is honest enough, and it tells you the category has not settled on a word.
The consequence for a buyer is that the name on the tin carries no information. You cannot shortlist on terminology, and you cannot shortlist on the service description either, because every provider in the category describes roughly the same set of activities: technical fixes, entity work, answer-shaped content, third-party corroboration and tracking. What actually differs between two providers is the measurement discipline underneath, and that is invisible from a homepage.
So the evaluation has to be run on the method rather than the offer. Everything below is a check a buyer can run in an afternoon, without our help and without ours being the answer.
Look at the measurement method, the engines covered, and whether the provider publishes case studies with dates attached. Provider A emphasises its platform, Provider B its senior-led delivery, Provider C its published research.
illustrative, not a capture An answer built from provider marketing can compare stated emphases. It cannot check whether a published number is reachable under the method that produced it, which is the one test that costs a buyer five minutes and separates the field. Written to show the shape of an engine answer. It has no session, no capture date and no sample size, so it is not evidence about any category, including yours.
Audit the arithmetic before you audit the agency
The strongest signal a provider gives you is a published number you can test against their own stated method. Most providers publish something: a ranking, an index, a benchmark, a client figure. Take the definition they give, take the number they printed, and check whether the second is reachable from the first.
Here is a worked example from the audited set, and we want to be precise about what is theirs and what is ours. aeoengine.ai publishes an Answer Index: four monthly editions launched between 2026-07-28 and 2026-07-31, each described on their own methodology page as 300 questions run across four engines, which is 1,200 answers per edition. That is their published method, stated by them. Their methodology also defines appearance rate as the share of the 300 tracked questions in which a firm is named. That definition is theirs too.
The arithmetic is ours, and anyone can repeat it. If appearance rate is a count of questions out of 300, its possible values are multiples of one three-hundredth, or roughly 0.333 percent. Rounded to one decimal place, the reachable low values are 0.3, 0.7, 1.0 and so on. The tail of their published National 100, positions 59 to 100, carries values of 0.9, 0.8, 0.6, 0.5, 0.4, 0.2 and 0.1 percent. Most of those cannot be produced by k divided by 300 under any rounding rule. Either the definition is not the calculation, or the tail is smoothed. We do not know which and we are not asserting intent. We are saying that a reader who spends five minutes with a calculator finds a gap between a published definition and a published number, and that is exactly the check to run on any provider, including on us.
A second, easier version of the same check is internal consistency. The same publisher's context columns disagree with themselves inside one month: one law firm is shown at a domain rating of 91 with 2.4 million organic sessions in one edition and a domain rating of 71 in another; a second firm is shown at 76 with 210,000 sessions in one and 0.3 with none in the other. Context columns that contradict each other in the same month are a sign that an entity-resolution step is quietly failing, and if it fails in the filler columns it is worth asking what it does in the columns that carry the ranking.
None of this makes them a bad provider. Parts of that index are better practice than most of the category manages: the question set is frozen and versioned, there is a public correction log, and retired sample prompts are published so a reader can see the shape of the questions without being able to game them. We would credit all three. The point is narrower. A provider who publishes enough for you to check them is doing something valuable even when the check fails, and a provider who publishes nothing checkable has not given you anything to evaluate.
Ask what the provider withholds
Every AI visibility number is a ratio computed over a set of prompts somebody chose, run some number of times, in some window. Drop the prompts you never win and the share rises. Add prompts containing your own brand name and it rises again. Restrict the competitive field to three named rivals rather than every company the answers mentioned and the number can double while not one captured answer has changed.
That is why the useful question is not what a provider reports but what they hold back. Ask for the noise floor: below what movement do they call a change run-to-run variance rather than progress. Ask whether any figure in the last quarter was withheld for failing that bar, and what it was. A measurement discipline that has never produced a held figure is not a discipline.
Our own answer, so this is not a demand we make of others only. We hold any citation figure that has not run at least three times. Every client citation share we have measured so far is a single-sample reading with roughly 25 points of margin, so none of it is on this site. We can show you a 488-point local grid, a page-by-page indexation audit and a controlled ranking experiment. We cannot show you a client citation percentage, and the reason is that it would not survive being re-run.
Ask about engine coverage the same way, and expect the honest answer to be short. We measure ChatGPT and Gemini through a vendor that renders those consumer surfaces, plus Google AI Overviews through search results with the overview expanded. We do not measure Claude at all, because there is no ground-truth path to the consumer product and a model-level API is a different surface with different retrieval behaviour. We do not measure Perplexity either, because its consumer product disallows automated access in its own robots file. Any provider whose coverage list includes every engine you have heard of is either scraping surfaces whose terms forbid it or counting an API as the consumer product.
Ask for the control group
The single question that separates an operator who measures from an operator who reports is whether they have ever run a change against an untouched cohort. Without one, every result is a movement that happened after something, and search results move on their own.
Our own example is a failure, which is the only kind we can publish without a client sign-off. In August 2026 we changed page titles on a video production company in the San Francisco Bay Area. The homepage average position went from 16.8 to 34.5 overnight, and homepage impressions went from 889 a day to 538. The reason we know it was us rather than the market is that 34 untouched pages on the same site moved 0.8 positions over the same period, which is nothing. The nine retitled pages moved 3.1. Search Console, day over day, 2026-08-10.
The mechanism was legible once we looked at term level. For one commercial query the homepage fell from position 5 to 43 while the corresponding service page rose from 36 to 15. Google had moved the term to a page that ranked worse for it. We rolled the titles back the same day.
A buyer should ask any provider for the equivalent story, and be suspicious of a portfolio with no negative results in it. Work that is measured produces failures. Work that only produces wins is either extraordinarily lucky or is not being measured against anything.
Scope, terms, and what a credible engagement contains
Five things belong in the scope of work, and we publish this list precisely so a buyer can hold us to it as well as anyone else. A measured baseline of your citation share across the engines the provider can actually reach, not an estimate derived from traffic. Named competitor benchmarking, so the number has context. Execution that ships, meaning the technical and content fixes are done rather than recommended. Re-measurement at the end against the same baseline, so the work is accountable to the metric it claims to move. And plain reporting with the methodology shown, with no multiples nobody can verify.
On terms, the accessible end of the market splits. aeoengine.ai publishes 90-day rolling terms, and Breaking B2B's pricing FAQ, re-fetched on 2026-08-16, says its retainers run six or twelve months and that it does not run shorter. Two of the six providers we audited publish a performance guarantee on their own site. Those are their published terms as of 2026-08-15, and we have not tested how either is applied in practice. We do not publish a guarantee. Our reasoning is that a guarantee on a probabilistic output is either hedged into meaninglessness or is a commitment the seller cannot control, and we would rather say that than invent one to match. A buyer is entitled to treat that as a mark against us.
Ask who owns what at the end. Ask whether the prompt set is yours and whether you can take the raw captured answers with you. Ask whether a strategist is on the account or whether the engagement is a platform your own team will be running, because the second is a real and reasonable model and it is priced differently. One provider we inventoried is explicit about it: their site is 65 pages of product documentation and 58 essays against three blog posts, which is the shape of a company selling software, and they publish seven legal pages including a master agreement, a data processing addendum and a subprocessor list. If your procurement team needs that furniture, that matters, and we do not have it.
Finally, ask what would make them tell you not to buy. Our version of the answer is that the free assessment exists partly to produce that conclusion: if AI answers do not name businesses in your category yet, the measurement will show it, and the correct outcome is that you spend nothing. A diagnostic seller who cannot describe the circumstances in which they would walk away is not selling a diagnosis.
Questions, answered plainly.
Is an answer engine optimization agency different from a GEO or LLM SEO agency?
No. The four terms name one service and the split exists because buyers search different phrases. One provider we inventoried in August 2026 runs three separate service URLs for a single AI-search offering, one per search term. Shortlist on measurement method rather than terminology, because the terminology carries no information.
What should I ask an AEO agency on the first call?
What is the prompt set and can I see it. How many times is each prompt run. What is the window, stated as dates. Which engines do you cover, and which can you not reach and why. Below what movement do you call a change noise. And: name a figure you withheld in the last quarter and why. A method that never produces a held figure is not a method.
How do I check a provider's published numbers?
Take their stated definition and test whether the printed value is reachable from it. One provider defines appearance rate as a share of 300 tracked questions, which makes the possible values multiples of roughly 0.333 percent, yet publishes tail values of 0.9, 0.8, 0.6, 0.5, 0.4, 0.2 and 0.1 percent. The definition and the numbers are theirs; the arithmetic takes five minutes and anyone can repeat it.
Why does a control group matter for AI visibility work?
Because search results and AI answers move on their own, so any change measured after an edit is a coincidence until an untouched cohort rules that out. When we broke a client's homepage ranking with a title change in August 2026, the evidence it was us rather than the market was that 34 untouched pages on the same site moved 0.8 positions while the homepage moved from 16.8 to 34.5 overnight.
Should an AEO agency guarantee results?
Two of the six providers we audited in August 2026 publish a performance guarantee on their own site, and we have not tested how either is applied. We do not publish one, because a guarantee attached to a probabilistic output is either hedged until it means nothing or promises something the seller does not control. That is a defensible reason and it is also a real disadvantage for us against providers who do offer one.
What is measured, and what this page is not.
This is an explainer. It carries no figures, and it is not a reading of your category. The disclosure below states the instrument that produces the numbers the essay refers to, so the distinction is on the page rather than assumed.
- instrument
- Caul
- what was measured
- Nothing on this page. Where the essay refers to citation share, that figure is produced separately, per account.
- how
- A prompt set written once for a category and then frozen, run against every engine in clean sessions, with each answer stored unmodified.
- over what window
- Reviewed on 2026-08-16. The engines change, so read the essay against the date on the byline.
- what this cannot tell you
- An explainer is not evidence about your category. Being named is not being recommended, and it is not traffic or revenue. Any figure about your own visibility has to come from a capture of your own category, carrying its sample size and its window.
The rest of the cluster.
See where you stand.
The audit is a real sweep of your category, benchmarked against competitors you name, delivered on a call so the findings get explained rather than emailed. You keep the report and the underlying data whatever you decide afterwards.