Entity clarity: why AI does not know who you are
Entity clarity is whether an answer engine can resolve your brand name to one specific organisation rather than to a word it already knows, and a coined name buys that clarity only after independent sources corroborate it, never on the day you register it.
Key takeaways
- Most brands that say an engine does not know them have a resolution problem, not an absence problem. The engine has plenty of material for the string and cannot decide which thing the string refers to.
- A coined one-word name is not protected by being unique. An engine's first move on an unfamiliar string is to normalise it toward the nearest word it already has evidence for.
- sameAs is an identity assertion, not a citation lever. Google's own guidance says structured data is not required for generative AI search, and that is exactly the claim sameAs was never making.
- Corroboration is measured across source types, not source counts. Ten directory profiles are one type repeated ten times, and one sentence repeated identically across four types beats a fuller description phrased four different ways.
- Entity work keys off a domain, because every corroborating source embeds a domain string. Freeze the domain and the one-sentence description before the first profile exists, or you buy the job twice.
What an engine is doing when it does not know you
The complaint usually arrives in the same words. We asked ChatGPT about our company and it described a different business, or it described nothing at all. The natural reading is that the engine has never heard of you, and the natural prescription that follows is to publish more, so that it does. That prescription is right about a minority of cases and wrong about most of them, and the difference is worth more than any content plan.
In most cases the engine has abundant material for the string you typed. What it does not have is a way to decide which thing the string refers to. That is a resolution failure rather than an absence, and the two look identical from the outside: you asked, you were not described, you concluded you were invisible. Underneath they are opposite conditions with opposite fixes. Absence means you are not in the retrievable pool for the question at all, which is an authority problem and is slow and expensive to solve. Contested resolution means the pool contains several unrelated things wearing your name, which is a disambiguation problem and is comparatively cheap.
Getting the diagnosis backwards is expensive in a specific direction. Publishing volume against a contested token adds mass to a string that is already carrying several referents, and the reasoning here is ours rather than a measurement: more documents using an ambiguous name give the engine more evidence that the name is ambiguous. The fix for contested resolution is more material that ties the name to one referent, in places you do not own, rather than more material about you.
It is easier to show this than to argue it. On 2026-08-16 we asked DataForSEO's ChatGPT scraper, with web search enabled, a single question about our own brand: who or what is Caldrin. The answer named four referents. A fictional satyr illusionist in a tabletop setting. An ascended deity of civilization, peace and craftsmanship in a published worldbuilding wiki. An uncommon given name whose claimed etymologies it flagged as user-submitted rather than documented. And a historical pharmaceutical trade name. It cited four sources: a worldbuilding wiki twice, a baby-name site, and a drug-name aggregator. It named no company, ours or anyone else's.
Its first sentence was the diagnosis: “Caldrin doesn’t appear to have one universally recognized meaning or referent.” Its last move was better. It asked us which Caldrin we meant. An engine that closes by requesting disambiguation has told you precisely what your entity problem is, in the one place most people never think to look, which is the answer itself rather than a dashboard reading of it. Two limits belong in the same breath. This was one run, one engine, one day, and answers vary run to run, so under our own bar it illustrates a failure mode and measures nothing. And the pharmaceutical claim is the engine's, sourced to an aggregator we have not checked, which is itself instructive about the quality of source an engine will reach for when better corroboration does not exist.
The nine results that belong to somebody else
The same day, one run each, we pulled Google's organic results for the bare token in the United States on desktop, through DataForSEO. Nine organic results came back on the first page. A rock band from West Lothian in Scotland held three of them, across a Facebook page, a Bandcamp profile and a Spotify album. A DeviantArt user in Ohio, active for twenty-two years, held one. The ascended deity from the worldbuilding wiki held another. The rest went to an AI-music profile, a Steam Community account, a tennis player listed on the Davis Cup site as playing for Jamaica, and an author page on Amazon.
None of the nine was a company. That is worth pausing on, because it is not the result the theory predicts. The token is rare enough that no established business holds it, and rather than sitting empty it has been taken up by individuals and by fiction. A worldbuilding wiki is not weak evidence to a retrieval system. It is a structured page, with a subject, a description, a category and a date, describing one named thing consistently, and it has been sitting there since 2024. The Scottish band's three results are the same phenomenon in its healthy form: a primary presence, platform listings, and a catalogue, all describing one act in the same terms. Nobody involved was doing entity work. They were making music and filling in profile fields, which turns out to be the same activity.
A Google AI Overview block also fired at the top of that page. We did not expand it, so we do not know what it said, and a block firing tells you an overview was present and nothing about its contents. The distinction matters enough that it has its own essay in this set. It is recorded here because leaving it out would be tidier and less true.
The second query is where the argument actually lives. We added the category word and asked for “caldrin agency”. Nine organic results came back and every single one of them was a Cauldron. A domain called cauldronagency.com. A marketing shop called The Digital Cauldron, twice, once on its own domain and once on a directory profile. A technology-led entertainment studio called The Cauldron Company. A vegetarian food brand called Cauldron Foods, via a trade-press story about its social agency appointment. A non-profit theatre called Creative Cauldron. A Wikipedia category for games made by a developer called Cauldron. A recruiting-software company called Cauldron on Crunchbase. And a web development shop at cauldron.re, which sells, among other things, SEO.
Zero results for the string we typed. Google read our coined token as a misspelling of a common English word and served the word instead. This is the finding that should change how the category talks about naming, because it inverts the usual advice at exactly the moment the advice is being taken. Uniqueness in a dictionary is not uniqueness in an index. An unfamiliar string does not arrive at a retrieval system as a protected namespace, it arrives as a probable typo, and the system's first and most reliable move is to normalise it toward something it already has millions of documents about.
Why a coined name is an endgame, not a starting position
The orthodox case for a coined one-word brand is genuinely strong and we accepted it. A made-up token has no dictionary meaning to compete with, no partial-match dilution, and no category ceiling if you later change what you sell. Brand search accrues cleanly to one string, because nobody types a coined word by accident. When somebody says the name to an engine, there is only one thing they can mean. Our own internal decision document says close to that, in close to those words, on the day the name was chosen.
Every one of those properties describes a coined brand that has already been corroborated. None of them describes a coined brand on the day it is registered, and the sleight of hand in the orthodoxy is that it presents an outcome as an input. Zero disambiguation cost is not a property of the string. It is a property of a string that enough independent sources have attached to one referent, and it is purchased, over months, with the corroboration work the name was supposed to make unnecessary. Before that purchase clears, a coined token has the opposite profile: no prior, no category attachment, and a spelling neighbour with a hundred years of documents behind it.
A descriptive name inverts every term of that trade. It is diluted from the first day and stays diluted, it competes with every firm that describes itself the same way, and it will not carry a distinctive entity of its own without significant work. What it does buy is immediate legibility. An engine that has never encountered your company can still place a descriptive name in the right category on first contact, because the words are doing the resolution that corroboration has not yet done. That is a real advantage and it is an advantage about time, not about quality. A coined name is a better asset and a worse starting position, and anybody choosing one should be told which of those they are buying first.
There is a second company on our token, and it is instructive that we found it only by asking for it. An exact-phrase query on 2026-08-16 returns Caldrin Systems at position one on its own name, at caldrinsystems.com. Their page title reads “AI agents for high-friction workflows”, and their own description, as Google renders it, says they build deterministic AI for industries where every dollar must reconcile, in Australian custom fabrication and Australian private markets.
A collider in a distant category is easy to separate, and the intuition most people have is that this is one of those. It is not, and the reason is the part worth taking away. Compress both companies to the single sentence an engine has room for and both of them are an AI company. Their industries diverge, their buyers do not overlap, and none of that helps, because the compression happens before the divergence is reached. Semantic adjacency raises the disambiguation load rather than lowering it. It also raises the standard for how we describe ourselves, since a description that could be true of them is not doing any resolving.
They hold position one for their own name and we hold nothing, and that outcome is correct. Their domain resolves and serves a site. On 2026-08-16 ours had nameservers at Cloudflare and no A record at all. A retrieval system cannot resolve a name to an organisation when the organisation has no address to resolve to, and no amount of markup on a page that is not being served will change that.
What sameAs is for, and the claim it cannot carry
sameAs is an array on an Organization object in your structured data. It lists other pages on the web that describe the same entity: a Wikidata item, a Crunchbase record, a LinkedIn company page, a licensing-body listing. It is one of the most widely recommended tactics in this category and it is routinely sold as a way to get cited by AI. It is not that, and the confusion is doing real damage to how the work is scoped and priced.
Google's published guidance on optimizing for its generative AI features says, under a heading warning against overfocusing on structured data, that structured data is not required for generative AI search and that there is no special schema.org markup you need to add, while recommending you keep using it for the rich results it does make you eligible for. That is at developers.google.com/search/docs/fundamentals/ai-optimization-guide, retrieved 2026-08-16. Read it precisely, because the sentence is doing two things. It closes the argument that schema is a citation lever. It says nothing at all about entity resolution, which is the job schema is genuinely doing and the only job sameAs was ever performing.
So the honest framing is narrower and more useful than the one being sold. A sameAs array is not a request to be cited. It is a claim about identity, addressed to a system that is trying to work out whether the Caldrin on this page is the Caldrin on that one. Different claim, different mechanism, and it should be priced and reported as entity hygiene rather than as a visibility lever. We ship it on every account. We do not bill it as a citation driver, and we have no controlled test showing it moves citation rate. We are not aware of a published one from anybody, ourselves included, which is worth saying plainly given how confidently the tactic gets sold.
The deeper constraint is that your own markup is a self-assertion, and self-assertion is structurally discounted. This is not a scoring preference that could be tuned away next quarter, it is an architectural limit. There is no registry of true things. No crawler can telephone a county clerk to confirm you exist. At web scale, a system that answers questions by synthesising sources has no option but to treat agreement across independent sources as a stand-in for truth, which means consistency, not truth, is the quantity actually being measured. Your own site is the baseline, the sentence the other sources have to agree with. It is never the evidence.
That reframes the work from a checklist into something more awkward and more effective. What matters is the diversity of source types, not the count. A primary site you control. A reference-style third-party entry written in a neutral register. Independent editorial that describes you without your involvement. And an established third party that already has the engine's trust naming you and linking to you. Ten directory profiles look like ten sources and function as one type repeated ten times. One of each of the four beats thirty of the first.
Consistency then beats completeness, which is the rule most real companies break without noticing. A true and detailed fact stated four different ways across four sources reads to the machine as unclear. A modest fact stated identically everywhere reads as established. Same legal name, same one-sentence description, same category, same founder names, everywhere, including the places that feel too small to matter. Real businesses tend to have the substance and none of the signature, and closing that gap requires no invention at all, only transcription. One caution to end on: nobody can promise you a knowledge panel. You can build every input to one and Google can still decline. A provider selling you the panel itself is selling an output they do not control.
We are inside this problem, not describing it from outside
Everything above is a report from a company currently failing at it. We chose a coined token knowing the theory, and the captures in this essay are our own brand returning a rock band, a fictional god, and eight companies with a different spelling. Nothing here is a recovery story, because there has been no recovery yet.
The worse admission is a process failure rather than a naming one. We shipped a build whose canonical URL named a domain we had never registered. Not lapsed, not expired, never bought. It sat in our own strategy document as the decided primary domain, was read by several people over several weeks, and was carried into code as a canonical string, and nobody ran a registry query against it. When somebody finally did, on 2026-07-17, the registry returned no match. It still returns no match today, on 2026-08-16. The domain is not named on this page on purpose, because publishing an unregistered lookalike string is an invitation.
The detail that makes it a genuinely useful failure is how it hid. The same document recorded that domain as measured at zero organic today. Read quickly, that is an ordinary line about an early-stage property with no traffic yet. Read correctly, it meant the hostname did not exist and the instrument had returned an empty result. An empty return had been read as a low value rather than as a question about the instrument, and a measurement of nothing had quietly become a baseline. It is the same error as a percentage published without its denominator, and it is why we treat any absence as a claim that needs a second and third instrument before it is believed.
The operational rule that follows is a sequencing gate, and it is the most transferable thing in this essay. Entity work keys off a domain. The reference entry embeds a domain string, so does the press mention, so does every directory profile, so does the sameAs array itself. Build thirty corroborating sources and then change the domain and you have bought the job twice, and you have also manufactured the exact inconsistency the work existed to remove: two hostnames and two spellings across your sources is what unclear looks like from the inside. Freeze the domain, and freeze the one-sentence description that every source will carry, before the first profile is created. Neither is a design decision and both get treated as one.
Our own position, stated plainly, because a page arguing for published limits should carry its own. caldrin.co was registered on 2026-07-16, so it has no domain age and no history to inherit or to audit. On 2026-08-16 it had Cloudflare nameservers and no A record, which means the primary source type does not exist yet, and the primary source type is the one that gets discounted anyway. The three types that carry actual weight have not been started. We are at the beginning of a months-long clock, and the reason we can describe the mechanism accurately is not that we have beaten it.
If you want to know where you stand, the check takes about ten minutes and needs no tooling. Search your bare brand name and read what comes back, counting how many of the results are you. Search your brand name plus your category word, and watch for whether the engine quietly substitutes a more common word. Then ask an engine with web search enabled who you are, and read the answer for a clarifying question, because an engine asking which one you mean has diagnosed you for free. If more than one referent comes back, entity work belongs in front of your content plan rather than behind it. Then run all three again a week later, because one capture is one run, and a single answer from a system that does not repeat itself is a story rather than a reading.
Questions, answered plainly.
Why does AI describe my company as a different business?
Usually because the engine has abundant material for your name and no way to decide which thing the name refers to. That is a resolution failure rather than an absence, and the two look identical from the outside: you asked, you were not described, you concluded you were invisible. Absence means you are not in the retrievable pool at all, which is an authority problem. Contested resolution means the pool contains several unrelated things wearing your name, which is a disambiguation problem and is comparatively cheap to fix.
Does a sameAs array help me get cited by AI?
No. A sameAs array is a claim about identity, not a request to be cited. Google's published guidance on its generative AI features states that structured data is not required for generative AI search and that there is no special schema.org markup you need to add, while still recommending it for the rich results it makes you eligible for. Ship sameAs as entity hygiene, price it as entity hygiene, and treat any provider billing it as a citation driver as having skipped that sentence.
Is a coined brand name better for AI visibility than a descriptive one?
Eventually yes, and on day one no. A coined token has no dictionary competition and no category ceiling, but those are properties of a coined name that independent sources have already attached to one referent, not properties of the string itself. Before that corroboration exists, an engine's first move on an unfamiliar string is to normalise it toward the nearest word it already knows. Our own coined name plus a category word returned nine organic results on 2026-08-16 and every one of them was a differently spelled common word.
How many sources do I need before an engine resolves my name?
The count is the wrong unit. What is measured is diversity of source types: a primary site, a neutral reference-style entry, independent editorial, and an established third party naming you. Ten directory profiles are one type repeated ten times, and one of each of the four beats thirty of the first. Consistency then beats completeness, because a true fact stated four different ways reads as unclear while a modest fact stated identically everywhere reads as established. No number guarantees an outcome, and nobody can promise you a knowledge panel.
What is measured, and what this page is not.
This is an explainer. It carries no figures, and it is not a reading of your category. The disclosure below states the instrument that produces the numbers the essay refers to, so the distinction is on the page rather than assumed.
- instrument
- Caul
- what was measured
- Nothing on this page. Where the essay refers to citation share, that figure is produced separately, per account.
- how
- A prompt set written once for a category and then frozen, run against every engine in clean sessions, with each answer stored unmodified.
- over what window
- Reviewed on 2026-08-16. The engines change, so read the essay against the date on the byline.
- what this cannot tell you
- An explainer is not evidence about your category. Being named is not being recommended, and it is not traffic or revenue. Any figure about your own visibility has to come from a capture of your own category, carrying its sample size and its window.
The rest of the cluster.
See where you stand.
The audit is a real sweep of your category, benchmarked against competitors you name, delivered on a call so the findings get explained rather than emailed. You keep the report and the underlying data whatever you decide afterwards.