Skip to content
Home-services AI measurement

How We Measured Home-Services Visibility in AI Search

An AI-visibility score without its prompt list, surfaces, geography, date and failure log is impossible to interpret. A clean zero can hide excluded errors. A familiar source can look important simply because it appeared often. A blended score can merge two systems that produced very different answers.

The useful question is narrower: what did a named brand do inside a fixed, repeatable observation set?

Caldrin tested that question with one home-services panel. The panel is an example of measurement discipline, not an industry benchmark. It cannot tell us how every homeowner searches, how every AI product behaves or whether visibility creates revenue.

Measure mentions, citations and recommendations separately

AI-search visibility has at least three observable states, and they should not be treated as synonyms.

Measure mentions, citations and recommendations separately table
StateOperational definitionWhat it does not establish
MentionThe answer names the brand.The answer does not necessarily link to the brand, endorse it or send a lead.
CitationThe answer attributes a source through a linked or recorded citation. A brand-owned citation points to the brand's own domain.A citation is not automatically a recommendation, a trust score or a commercial outcome.
RecommendationThe answer explicitly shortlists, endorses or orders an entity.Machine-extracted names alone do not prove recommendation order or sentiment.

Caldrin's panel stored mentions and owned-domain citations as separate fields. Recommendation ordering was not reliable enough for headline use, so it stays NOT_MEASURED in this page.

The measured panel used one frozen cohort and two surfaces

Caldrin's internal panel used 20 frozen prompts: 10 about foundation repair and 10 about HVAC. Each prompt was measured on ChatGPT consumer ground truth and rendered Google AI Overviews. The requested and effective geography for eligible observations was the United States at country grain.

Foundation repair is one half of the prompt cohort here. The foundation-repair search system owns the category's service, market, profile and lead-measurement decisions.

Collection ran sequentially in one UTC window, from 2026-08-31T14:33:27.406Z to 2026-08-31T17:51:13.325Z. Every one of the 40 prompt-and-surface cells reached 30 eligible observations. The panel retained 1,263 attempts in total: 1,200 eligible observations, 57 failures and 6 quarantines.

Caldrin internal panel: 20 frozen prompts, two measured consumer-facing surfaces, United States at country grain, collected in one same-day UTC window. The matrix does not represent every query, model, Google AI surface, local market or persistence across time.
ClusterFrozen promptsMeasured surfacesEffective geography for eligible observationsCollection window
Foundation repair10ChatGPT consumer ground truth; rendered Google AI OverviewsUnited States, country grain2026-08-31T14:33:27.406Z to 2026-08-31T17:51:13.325Z
HVAC10ChatGPT consumer ground truth; rendered Google AI OverviewsUnited States, country grain2026-08-31T14:33:27.406Z to 2026-08-31T17:51:13.325Z

Panel scope: Caldrin internal panel, collected in one UTC window from 2026-08-31T14:33:27.406Z to 2026-08-31T17:51:13.325Z: 20 frozen prompts (10 foundation repair, 10 HVAC), ChatGPT consumer ground truth and rendered Google AI Overviews, United States at country grain; 1,200 eligible observations from 1,263 attempts, with 57 failed and 6 quarantined. The panel does not represent every prompt, model, Google AI surface, date or local market, and it does not establish persistence across time.

Gemini and model APIs were outside this measured cohort. Their exclusion tells us nothing about Caldrin's performance on those surfaces.

Eligible observations are different from failed and quarantined attempts

An eligible observation was a successful, grounded output under the panel's inherited rules. A failed attempt never became a negative brand observation. A quarantined attempt stayed in the operational record but did not enter the eligible denominator.

Caldrin frozen 20-prompt US-country panel, collected from 2026-08-31T14:33:27.406Z to 2026-08-31T17:51:13.325Z. Failures and quarantines are retained run states, not negative visibility observations.
SurfaceAttemptsEligible observationsFailedQuarantined
ChatGPT consumer ground truth617600170
Rendered Google AI Overviews646600406
Total1,2631,200576

All 40 prompt-and-surface cells reached the target of 30 eligible observations. Individual cells required 30 to 35 attempts under a maximum of 40.

Accounting note: Caldrin internal panel, collected from 2026-08-31T14:33:27.406Z to 2026-08-31T17:51:13.325Z, across 20 frozen prompts, ChatGPT consumer ground truth and rendered Google AI Overviews, United States at country grain. The table separates 1,200 eligible observations from 57 failed and 6 quarantined attempts. Failures and quarantines are retained run states, not evidence that a brand was absent.

Caldrin did not appear inside this panel

Caldrin was not named and caldrin.co was not cited in any of the 1,200 eligible observations. No prompt in the frozen 20-prompt cohort produced an eligible Caldrin appearance.

Bounded result: 0/1,200 eligible observations and 0/20 frozen prompts with an eligible Caldrin appearance. The 1,200 eligible observations came from 1,263 attempts, with 57 failed and 6 quarantined, across ChatGPT consumer ground truth and rendered Google AI Overviews at United States country grain from 2026-08-31T14:33:27.406Z to 2026-08-31T17:51:13.325Z. This result does not prove that Caldrin is absent from other prompts, surfaces, dates, countries or local markets.

Caldrin's citation rate is NOT_MEASURED, not 0%. The rate is conditional on an eligible Caldrin appearance, and that denominator was empty.

This is the uncomfortable value of a baseline: it records the result that occurred, without polishing it into a better story.

Bounded panel observation

0 Caldrin names or owned-domain citations / 1,200 eligible observations

0 / 20 frozen prompts with an eligible Caldrin appearance

Caldrin citation rate: NOT_MEASURED

Frozen 20-prompt panel across ChatGPT consumer ground truth and rendered Google AI Overviews, United States at country grain, collected from 2026-08-31T14:33:27.406Z to 2026-08-31T17:51:13.325Z. The result does not establish absence outside this panel.

Source mapping shows recurrence, not trust

Eligible outputs in the panel cited 783 distinct domains. The source taxonomy was deliberately coarse: 543 of those 783 distinct domains were classified OTHER by the bounded rules. OTHER means unclassified by those rules. It does not mean low quality.

Observed recurrence in Caldrin's frozen 20-prompt US-country panel, collected from 2026-08-31T14:33:27.406Z to 2026-08-31T17:51:13.325Z. Recurrence is not trust, quality, authority, accuracy, recommendation rank or market share.
SurfaceDomainObservations citing the domainEligible surface denominator
ChatGPT consumer ground truthdevelopers.google.com168600
ChatGPT consumer ground truthsupport.google.com166600
Rendered Google AI Overviewsyoutube.com108600

Recurrence note: Caldrin's frozen 20-prompt US-country panel covered ChatGPT consumer ground truth and rendered Google AI Overviews from 2026-08-31T14:33:27.406Z to 2026-08-31T17:51:13.325Z. It retained 1,200 eligible observations from 1,263 attempts, with 57 failed and 6 quarantined; recurrence was counted separately among 600 eligible observations on each named surface. Recurrence is not trust, authority, accuracy, recommendation rank or market share. One sample URL from each listed domain returned HTTP 200 on 31 August 2026; that check established reachability only, not passage quality or the source's role in an answer.

Rendered Google AI Overviews contained citations in 312 of 600 eligible observations in this panel.

AIO denominator: 312/600 refers only to eligible rendered Google AI Overview observations in Caldrin's frozen 20-prompt US-country panel across ChatGPT consumer ground truth and rendered Google AI Overviews from 2026-08-31T14:33:27.406Z to 2026-08-31T17:51:13.325Z. The full panel retained 1,200 eligible observations from 1,263 attempts, with 57 failed and 6 quarantined. 312/600 is not the frequency of AI Overviews across Google searches and does not describe other Google AI surfaces.

Rendered Google AI Overviews

312 citation-bearing observations from 600 eligible rendered Google AI Overview observations

312/600 eligible rendered-AIO observations in Caldrin's frozen 20-prompt US-country panel on 31 August 2026. This is not the incidence of AI Overviews across Google searches and does not cover other Google AI surfaces.

Repetition measures variation inside a window

One answer per prompt cannot show whether the answer or source set changes on the next attempt. This panel repeated each prompt-and-surface cell until it reached 30 eligible observations, then stored per-cell answer and source-overlap measures.

Those repetitions describe variation inside one short collection window. They do not establish stability or persistence across days, weeks, later reruns, changed prompts, changed locations or provider changes. A later baseline should preserve the cohort and measurement rules if the goal is a like-for-like comparison.

Variance boundary: The variance artifact covers all 40 completed prompt-and-surface cells in Caldrin's frozen 20-prompt US-country panel across ChatGPT consumer ground truth and rendered Google AI Overviews. Collection ran from 2026-08-31T14:33:27.406Z to 2026-08-31T17:51:13.325Z, retaining 1,200 eligible observations from 1,263 attempts, with 57 failed and 6 quarantined. The measures describe same-window variation only. Machine extraction does not establish recommendation order, sentiment or market share.

Organic and AI-source overlap is a co-occurrence check

Ten domains appeared in both the panel's 783-domain AI-source set and a supplied 16-domain organic comparison set. The union contained 789 domains.

The AI-source set comes from Caldrin's frozen 20-prompt US-country panel across ChatGPT consumer ground truth and rendered Google AI Overviews, collected from 2026-08-31T14:33:27.406Z to 2026-08-31T17:51:13.325Z. Domain-level co-occurrence only, not same-URL overlap, link intersection, causation or proof that organic rank caused AI citation.
AI-cited domainsOrganic comparison domainsDomains in both setsUnion
7831610789

The ten co-occurring domains were birdeye.com, firstpagesage.com, getcourtyard.ai, hookagency.com, intleacht.ai, plumberseo.net, sequoiageo.com, shiftflow.app, silverbackstrategies.com and webfx.com.

Overlap boundary: The AI-source set comes from Caldrin's frozen 20-prompt US-country panel across ChatGPT consumer ground truth and rendered Google AI Overviews, collected from 2026-08-31T14:33:27.406Z to 2026-08-31T17:51:13.325Z, with 1,200 eligible observations from 1,263 attempts, 57 failed and 6 quarantined. The supplied organic set is a separate dataset. 10/789 is domain-level co-occurrence, not same-URL overlap, a link intersection, causation or proof that organic ranking produced an AI citation.

Build a baseline that can survive a rerun

A useful AI-visibility baseline should make its choices visible before the results arrive.

  1. Freeze the prompt text and record why each prompt belongs in the cohort.
  2. Record the consumer-facing surface separately. Do not blend a chat interface, rendered search result and model API into one row.
  3. Store requested and effective geography. Country-grain evidence cannot answer a city or local-pack question. Use the multi-location SEO operating model for location-page, review and Business Profile operations.
  4. Define eligibility before collection. Keep failed and quarantined attempts in the operational ledger.
  5. Repeat each prompt-and-surface cell to a declared target and retain the attempt cap.
  6. Store mentions, owned-domain citations and recommendations as different fields.
  7. Map source recurrence with a surface denominator. Keep source quality as a separate review.
  8. Version the cohort. A changed prompt list creates a new measurement population.

The panel's rounded recorded actual/estimated ledger total was $3.2324 for this run. That amount is not exact, an invoice, a quote, a guaranteed rerun cost, a public rate or a cost benchmark.

Baseline scope: The process above is illustrated by Caldrin's frozen 20-prompt US-country panel across ChatGPT consumer ground truth and rendered Google AI Overviews, collected from 2026-08-31T14:33:27.406Z to 2026-08-31T17:51:13.325Z. It produced 1,200 eligible observations from 1,263 attempts, with 57 failed and 6 quarantined. The process defines an observation set. It does not promise more citations, rankings, leads or revenue.

Intervene, then remeasure the same population

An intervention log should name the changed page or asset, the change date, the intended measurement, the owner and the rollback state. The next run should reuse the frozen prompts, surfaces, geography, eligibility rule, repetition target and attempt cap.

A post-change difference is an observation, not automatic proof that the intervention caused it. Provider behaviour, answer variance, changed source availability and time can all move between runs. If the population changes, version it and stop calling the comparison like for like.

The rerun records what moved. It still does not tell you why. Causal credit needs a design that can rule out competing explanations.

What this panel cannot tell a home-services team

The panel does not support claims about:

  • universal AI-search visibility;
  • city, service-area, near-me or local-pack performance;
  • Gemini, model APIs or every Google AI experience;
  • which recurring source is most trusted, accurate or authoritative;
  • recommendation order, sentiment or competitor market share;
  • leads, pipeline, revenue, conversion or return on investment;
  • whether reviews, links, schema, content or domain consolidation caused a citation; or
  • persistence across days, weeks, changed prompts, changed locations or provider changes.

The page format is also an editorial direction. No counted live-SERP structural study was supplied, so dominant page type, competitor word-count range, entity consensus and NLP salience remain NOT_MEASURED.

Get the free AI Visibility Report

The AI Visibility Report is free. A short qualifying form comes first. If the account fits, Caldrin defines the prompt set, surfaces, geography, eligibility rules and rerun conditions before presenting the findings live.

The report does not promise increased citations, rankings, leads, revenue or inaccessible-surface coverage. It gives the team a repeatable observation set to inspect and remeasure.

If the blocked decision is broader than AI measurement, use the home-services search partner model to identify the right operating scope.