What we got wrong
This is a running, dated ledger of the measurements and changes Caldrin got wrong, kept because a control group is the only thing that separates an agency which learns from an agency with an alibi, and it opens with a page title we changed on 31 July 2026 that sent a client homepage from 16.8 to 34.5 in average position overnight while 34 untouched control pages moved 0.8.
Key takeaways
- The number that made our worst change readable was not the fall. It was the 0.8 positions that 34 untouched pages moved on the same day, because that is what ruled out an algorithm update, a hosting fault and a link cleanup all at once.
- An agency with no control cohort cannot separate its own damage from the weather, so it will never report any. The same gap makes its wins unattributable, which is the part nobody mentions.
- Rolling a change back does not roll the rankings back. A re-crawl gets a page reassessed, not rewound, so the cost of making a change and the cost of undoing it are not the same size.
- A title edit is an instruction about which page owns a query. Removing two words moved one head term from a homepage at position 5 to a service page at position 15, and the site kept the relevance while losing the position.
- Code has a compiler and copy does not. Every error in this ledger that reached a page was caught by a claim-by-claim audit, not by a test, which is why the audit is scheduled rather than occasional.
The first entry, 31 July to 10 August 2026
On 31 July 2026 we rewrote the title tag on a client homepage and asked Google to re-crawl it the same day. Google did. The next morning that page's average position in Search Console had fallen from 16.8 to 34.5, and the number of times it appeared in search had roughly halved, from 889 a day to 538. The client is a video production company in the San Francisco Bay Area. They are unnamed on this page because public naming sign-off is still pending, which is a permission gap rather than a vagueness in the record: every figure below sits in their delivery log against the date it was measured.
The old title carried the words corporate and commercial. The replacement dropped both and added the company name, which is the ordinary and usually correct thing to do with a homepage title. It was a considered change rather than a careless one, and the reasoning behind it was sound. Those two facts sit together far more often than agency writing admits, and pretending otherwise is what makes most published post-mortems useless: they are written as though bad outcomes only follow obvious mistakes, so a reader learns nothing transferable.
We reported it to the client in writing on 10 August, before we knew whether the page would recover, and we put the title back to exactly what it had been. The order matters to everything else on this page. A failure disclosed after the recovery is a story. A failure disclosed while it is still costing somebody money is a measurement, and it is the only version that carries any information about how a provider behaves when a result goes against them.
One caution went into that report and belongs here too. Putting a title back does not oblige Google to put the rankings back. A re-crawl gets a page reassessed, not rewound, so the cost of making a change and the cost of undoing it are not symmetrical. We have published no recovery figure. We will not publish one from the reading taken two days later either, because two further changes were made to that homepage before the reading date at the client's request, and we said at the time that this would make the read less clean. A contaminated reading does not become clean by being the only one available.
Most providers in this category publish customer results and case studies. Provider A, Provider B and Provider C all highlight wins such as traffic growth, improved mention rates and new enterprise pipeline.
illustrative, not a capture An engine can only summarise what has been written down. A failed experiment nobody published is missing from the answer because it does not exist as a source, not because it did not happen, which is the same reason a category can look uniformly successful from the outside. Written to show the shape of an engine answer. It has no session, no capture date and no sample size, so it is not evidence about any category, including yours.
The control cohort is the entire experiment
The figure that turns the paragraphs above into evidence is not 34.5. It is 0.8. Measured on 1 August against the previous week, the 34 pages on that site we had never touched moved an average of 0.8 positions, which is nothing. The nine pages retitled on 31 July moved 3.1. The homepage on its own moved 18.4.
Those cohort figures and the 16.8 to 34.5 figure are two different pieces of arithmetic over two different windows, and they do not reconcile. The difference between 16.8 and 34.5 is 17.7, not 18.4. One is a day over day reading of a single page, the other a cohort average against a prior week. Printing both side by side without saying so lets a reader treat one event measured twice as two pieces of corroborating evidence. It is the smallest entry on this page and the one most likely to be repeated by somebody else, which is why it is here rather than in a footnote.
The argument the cohort supports is narrow and strong. An algorithm update would have moved the untouched 34. So would a hosting fault, so would seasonality, and so would the spam-link cleanup that had been filed the day before the change. All four of those explanations predict collateral movement across the site, and none of them survived contact with a cohort that did not move. What remained was the thing we had done.
So the standing claim, and it is meant to be arguable: a provider without a control cohort cannot separate its own damage from the weather, and will therefore never report any. This is not an accusation of dishonesty. Most of that reporting is entirely sincere. A sincere explanation produced by an instrument that cannot distinguish between two causes is still worth nothing, and core update is the most convenient sincere explanation available in this industry, because it is unfalsifiable at the account level and arrives on a schedule.
The corollary is the part that should worry a buyer more. If untouched pages are never measured, then a rise after a change is exactly as unproven as a fall. An agency that cannot detect its own harm has also given up the ability to demonstrate its own value, and it usually has not noticed, because only one of those two failures ever produces an awkward conversation.
A site does not rank, a page does
The mechanism turned out to be more useful than the loss. The two words removed from that title were corporate and commercial, and the query that fell furthest was corporate video production san francisco. On that query the homepage went from position 5 to position 43. On the same query, on the same site, in the same window, the client's own corporate video page rose from 36 to 15.
That pair is worth reading twice. The site did not stop being relevant for the term. Google reassigned the term from one page to another, and the page it moved to ranked worse than the page it left. Relevance was roughly preserved and position was destroyed, which is not a distinction most reporting is built to show, because most reporting aggregates to the domain and a domain-level view of that week shows a modest dip and nothing else.
The general form of the finding is that a title is not a description of a page. It is an instruction about which of your pages owns a query. Edit it and you are not adjusting a relevance score, you are voting in an internal contest, and the contest can be won by a page you did not intend to enter. Every optimisation that treats title tags as page-local, which is to say almost every automated title audit ever run, is blind to this class of outcome by construction.
The same day produced a second and quieter version of the same error. We had also added the city name to six service page titles that had never carried one. That put six of the client's own pages into competition with their homepage for the same city queries, which is the identical mechanism running against us in a different direction. Five of the six were reverted. The sixth was deliberately left alone, because it was the one page that had gained from the change and taking the city back out would probably have handed that gain away.
The process lesson is duller than the mechanism and matters more. We changed two classes of page in one deployment, which meant that when the homepage moved we had to do forensic work to establish which of the two changes was responsible. Running them a week apart would have cost seven days and answered the question for free. We now stage title work in single classes, and the reason that rule exists is written down here rather than presented as something we always did.
Errors in the instrument
A measurement business fails in two places: in the instrument, and in the copy that quotes it. The instrument errors are the ones we catch, because software can be regression-tested and a number can be re-derived from a raw file. Four are on the record.
Until 2 June 2026 our own leaderboard computed a union across all runs and labelled it visibility. A union counts the prompts where a brand was named at least once, so it can only rise: every additional run is another chance to be included and no chance to be excluded. One internal account read 75 percent on that union where the corrected per-run rate was 38.2 percent. Nobody had gamed anything. The metric simply answered a different question from the one its label claimed, and it flattered a little more every month it ran. The union is now called reach, the per-run rate is the headline, and both definitions travel with the sample size.
Until 4 June 2026 a bug in our AI Overview capture stored empty stubs that were indistinguishable in the database from a genuine absence of an overview. Every reading taken before that fix was an artefact rather than a low number, and 81 rows were purged from both databases rather than annotated. That distinction is the whole point of the entry: a broken instrument does not produce cautious data, it produces confident data about nothing, and the only safe response is deletion.
A third error concerned the basis rather than the capture. Reading current archive flags against historical runs produced a 48-point swing in one client's favour in a single day, with no answer anywhere in the system having changed. Metrics are now computed inside a frozen prompt-set cohort stamped at execution time, and the tool refuses to draw one trend line across a cohort change. A related finding sits underneath it: two prompt sets of the same size are not the same basis, and one client's version one and version two were both 55 prompts and shared only 17 of them.
The fourth is a finding we withdrew rather than a bug we fixed. We had reported a gap between what an engine retrieved for a client and what it produced without retrieval, and treated the gap as a fact about that brand. It was not. Two comparison brands run through the same test collapsed harder than the client did, which made the gap an artefact of the provider's answer format rather than a property of any of the three brands. The finding had already been written up internally when the comparison ran. It was withdrawn, and the run that killed it was the one nobody had asked for.
Errors in the copy, where there is no compiler
The second class is harder, because prose has no type checker and a wrong sentence renders exactly as cleanly as a right one. On 15 August 2026 we audited every claim intended for this site, tracing each figure back to a raw measurement file rather than to the document that last repeated it. It produced three corrections, and the audit is now scheduled rather than occasional.
The first was a competitor's client result sitting in our own internal design brief as though it were ours. It had arrived from that competitor's own industry page, been noted, and then been carried forward by ordinary document reuse until it read as a Caldrin outcome. Nothing was invented at any step. A number simply lost its provenance while being copied, which is how most fabricated marketing claims are actually produced. It is not on this page in any form, including as an example, because the failure mode is exactly that competitor proof migrates through examples.
Correction two was a bounded market claim with nothing behind it. The phrase of the five providers we looked at had reached 14 places on this site. It had entered as an illustrative example in an internal messaging document, been copied as a template, and become load-bearing without anyone ever writing down which five. A bounded claim with no roster is less honest than an unbounded one, not more, because it implies a rigour a reader cannot check. The correct construction, which is now the only permitted one, is that five of the six AEO and GEO providers we audited in August 2026 gate or charge for the initial assessment, and the exception is CrowdReply, which is self-serve software and has no assessment to gate.
The third was arithmetic. Re-deriving the local-pack figure for 603 Basement Solutions from a 488-point geo-grid, we read the pack rank column as a boolean and produced 0 percent presence. It is a rank, so the true figure is 10 of 488, or 2.0 percent. Separately, an internal note had been carrying 6 of 168 from an earlier and smaller cut of the same grid, which is roughly 3 percent. Same magnitude, different denominator, and no way for a reader to tell which was which. The wrong figure never left the building, but it was one document from doing so, and the rule it produced governs every number on this site: a percentage without its denominator and its field definition is not a measurement.
The figures we are holding this month
A ledger of errors is incomplete without the numbers we could publish and are not. Withholding is cheaper to fake than a failure, so it deserves the same specificity.
On the same 488-point geo-grid, 603 Basement Solutions held a local-pack position at 10 of 488 points on 9 August 2026 and 26 of 488 on 14 August 2026. We publish the baseline and not the movement. Five days is inside our own two-week settling rule, there is no control group, and the gain concentrates in a single keyword, which are three separate reasons the delta is not yet a result. The re-measure lands at 30 days. If it holds, it becomes the first outcome claim this company owns, and it will be published with the same denominator it is being withheld under.
What the baseline can carry is a comparison, because the grid measured everyone on it. One competitor holds a pack position at 38.1 percent of all 488 points and the next largest holds 21.7 percent. That is a before-picture and a demonstration of the method, and it is more useful to a prospective client than a delta would be, because it sizes the gap they are actually trying to close.
The AI citation share for every client we measure is held. Each of those figures is a single sample with roughly 25 points of margin, which is wide enough that the number would be repeated in rooms where the margin was not. The condition for release is a re-run at three or more passes, and the reason it is a hard gate rather than a preference is that a single-sample citation figure is precisely the sort of number that gets quoted back to us six months later as a baseline.
The largest thing being withheld is not a figure at all. There is no revenue-attributed win in this company's evidence yet. What exists is a measurement capability producing diagnostics that are unusually specific: a 488-point geo-grid, a page-by-page indexation audit checked against Search Console, and a controlled title experiment with a 34-page control cohort and a rollback. Writing around that gap with softer verbs would be the one failure this ledger could not survive, so it is stated instead.
The objection, which is a good one
The obvious response to a page like this is that a published failure is still marketing, and a provider clever enough to publish one is clever enough to select a flattering one. That is correct, and it is the right thing to be suspicious of. So here are the three tests we hold this page to, offered mainly so a reader can apply them to anybody else's version.
The first test is whether the client saw it before the internet did. The title rollback went to the client on 10 August in their delivery log, before we knew whether the page would recover and while the loss was still live. A failure that surfaces publicly only once it has stopped costing anything has been recycled rather than disclosed.
The second is whether it carries its control and its artefact. What makes the entry above usable is the 34 untouched pages, and what makes the numbers checkable is that they sit in a dated log with the raw Search Console pull behind them. An anecdote about a change that went badly, with no cohort and no artefact, is a piece of atmosphere. It costs nothing to write and proves nothing about the writer's instruments.
The third is cost. This entry describes a fall of roughly 18 positions overnight on a client's homepage, caused by us, disclosed before the outcome was known. A failure that turns out on inspection to be a strength in disguise belongs to a different genre, and readers can tell which one they are reading.
Across the six AEO and GEO providers we audited in August 2026, three of them inventoried page by page, we did not find a published negative result. That absence is not evidence that they never had one. It is evidence about what this category believes publishing is for, and the belief is that a case study is a sales asset rather than a piece of evidence. Those two things have opposite requirements: a sales asset needs to be favourable, and a piece of evidence needs to be checkable.
The rule this page runs on is the same one that governs every figure on this site. If a number moved, say what else moved. If nothing else moved, that is the finding. The next entry will be dated too.
Questions, answered plainly.
What exactly did the title change do?
On 31 July 2026 we rewrote the title tag on a client homepage and asked Google to re-crawl it the same day. The next morning that page's average position in Search Console had fallen from 16.8 to 34.5 and its impressions had roughly halved, from 889 a day to 538. The words removed were corporate and commercial, and on the query corporate video production san francisco the homepage went from position 5 to 43 while the client's own corporate video page rose from 36 to 15. Google reassigned the term to a different page on the same site, and that page ranked worse than the one it left.
How do you know it was the title and not a Google update?
Because of the control cohort. Measured on 1 August against the previous week, the 34 pages on that site we had never touched moved an average of 0.8 positions, which is nothing, while the nine retitled pages moved 3.1 and the homepage moved 18.4 on its own. An algorithm update, a hosting fault, seasonality and the spam-link cleanup filed the day before would all have dragged the untouched 34 down as well. None of them did, which leaves the change we made.
Did rolling the title back restore the rankings?
We have not published a recovery figure and we are not going to publish one from the reading taken two days after the rollback, because two further changes were made to that homepage before the reading date at the client's request, which we flagged at the time as making the read less clean. Putting a title back does not oblige Google to put rankings back in any case: a re-crawl gets a page reassessed, not rewound, so making a change and undoing it do not cost the same.
Why is the client not named?
Public naming sign-off is still pending, so we describe them as a video production company in the San Francisco Bay Area. That is a permission gap rather than vagueness in the record. Every figure on this page sits in their delivery log against the date it was measured, and the client received the whole account in writing on 10 August 2026, before we knew whether the page would recover.
Is publishing your failures just a different kind of marketing?
It can be, which is why we hold this page to three tests a reader can apply to anyone else's version. The first is whether the client saw it before the internet did, and here the account went to the client while the loss was still live. The second is whether it carries its control cohort and the dated artefact behind the numbers, because an anecdote with neither costs nothing to write. The third is cost, because a failure that turns out on inspection to be a strength in disguise belongs to a different genre and readers can tell which one they are reading.
What is measured, and what this page is not.
This is an explainer. It carries no figures, and it is not a reading of your category. The disclosure below states the instrument that produces the numbers the essay refers to, so the distinction is on the page rather than assumed.
- instrument
- Caul
- what was measured
- Nothing on this page. Where the essay refers to citation share, that figure is produced separately, per account.
- how
- A prompt set written once for a category and then frozen, run against every engine in clean sessions, with each answer stored unmodified.
- over what window
- Reviewed on 2026-08-16. The engines change, so read the essay against the date on the byline.
- what this cannot tell you
- An explainer is not evidence about your category. Being named is not being recommended, and it is not traffic or revenue. Any figure about your own visibility has to come from a capture of your own category, carrying its sample size and its window.
The rest of the cluster.
See where you stand.
The audit is a real sweep of your category, benchmarked against competitors you name, delivered on a call so the findings get explained rather than emailed. You keep the report and the underlying data whatever you decide afterwards.