AI Content Marketing
016 Getting Cited by AI Search 1,732 words · 8 min

Entity and Brand Signals: Becoming a Known Thing to a Model

Your last twelve blog posts did almost nothing for your AI visibility. I know that’s a rough opening, but I’ve now watched this pattern across enough content programmes that I’m confident about the mechanism. The posts were fine. Some of them ranked. What they didn’t do was make the model any more certain about who you are, and that certainty is the gate.

Here’s the thing that took me too long to understand. When an LLM decides whether to name you in an answer, it isn’t primarily asking “does this brand have relevant content?” It’s asking a much dumber, more mechanical question first: can I resolve this string of characters to a specific thing I have coherent information about? If the answer is fuzzy, you get dropped. Not penalised, not ranked lower. Dropped, because naming an entity you can’t resolve is how a model generates a hallucination, and the retrieval and grounding layers sitting in front of these systems are tuned hard against that.

So entity seo for brands isn’t a rebrand of technical SEO. It’s the work of becoming unambiguous.

What “resolvable” actually means in practice

Let me give you a real shape of the problem, using a composite of two clients I’ve worked with this year.

Company A is a UK B2B SaaS doing subscription billing analytics. Five years old, £4m ARR, 40 staff. They call themselves “Cadence” on the homepage, “Cadence Analytics” in the footer, “Cadence Software Ltd” on Companies House, “CadenceHQ” on X and GitHub, and “Cadence Billing” in their G2 listing. Their founder is credited as “Jen Marlowe” on the blog, “Jennifer Marlowe” on LinkedIn, and “J. Marlowe” in the two podcast appearances that actually get transcribed and indexed.

That’s five brand strings and three person strings for one company and one human. Every one of them is a separate candidate entity as far as an extraction pipeline is concerned. The disambiguation work that a model would need to do to merge them is real work, and it’s work that doesn’t happen reliably at the scale these systems operate.

Company B is boring by comparison. One name everywhere. Same twelve-word description in their LinkedIn About, their Crunchbase profile, their Organization schema, their Companies House SIC description and the boilerplate at the bottom of every press release since 2023. Two founders, full legal names, same spelling in all nine places they appear.

Company B gets cited in Perplexity and ChatGPT answers about their category. Company A, which has better content and more traffic, mostly doesn’t. When I tested 40 category-relevant prompts across ChatGPT search, Perplexity and Google AI Overviews in July, Company B appeared in 23. Company A appeared in 4, and in two of those it was named as “Cadence” with a link to the chip design company Cadence Design Systems. Which is the failure mode, right there, made visible: an ambiguous string resolved to the more famous entity.

The four surfaces that do the work

I’ve ended up with a fairly fixed order of operations on this, mostly because I’ve tried it in other orders and wasted time.

1. Pick one name and stop being creative. One canonical brand string, used in that exact form everywhere. Including the awkward places: your invoice template, your job ads, your conference badge copy, the alt text you were going to use as a keyword slot. Legal entity name can differ (it usually must) but it should appear once, on your about page, explicitly connected: “Cadence is the trading name of Cadence Software Ltd, registered in England and Wales, company number 09876543.” That sentence is doing more for your resolvability than a month of posts.

Same discipline for people. Pick the form of each name you’ll use, then use it in the byline, the author bio, the LinkedIn headline, the podcast intro you send to hosts, the speaker bio you give to conference organisers. “Jennifer Marlowe” or “Jen Marlowe”, I don’t care, but pick.

2. Build one about-surface worth resolving against. Not a mission statement. A dense, factual page that a retrieval system can lift clean sentences out of. The pattern that works:

  • What the company does, in one sentence, no metaphors
  • Founded date, founding location, founders by full name
  • Headcount band, funding position or profitability status, ownership
  • The three to five things you actually sell, named as products
  • Named customers or sectors served, with numbers where you can
  • Named leadership, with roles and one credential each
  • Press mentions, awards, memberships, certifications, with dates

Company B’s about page is 700 words and reads a bit like a Wikipedia stub. That’s the point. I’ve seen LLM answers quote that page nearly verbatim, because it’s the easiest place in their whole web presence to extract a defensible claim from.

3. Get corroborated by things that aren’t you. This is where most programmes underinvest, and it’s the highest-leverage surface. A model’s confidence in an entity goes up when the same facts appear in sources that don’t share a domain. Self-description is weak evidence. Third-party agreement is strong evidence.

Concretely, and roughly in order of effort-to-return for a UK company:

SurfaceRough effortWhy it matters
Companies House (description, officers, filing currency)1 hourAuthoritative, heavily crawled, free
LinkedIn company page + all staff profiles aligned3 hoursExtraction-friendly, high crawl frequency
Crunchbase profile, claimed and complete2 hoursFeeds a lot of downstream datasets
G2 / Capterra / Trustpilot, category and description consistent3 hoursCategory association, review text as corroboration
Wikidata item (not Wikipedia)2–4 hoursExplicit machine-readable identity graph
3–5 podcast appearances with published transcripts10 hoursNamed-person corroboration, quotable text
Trade press quotes with your exact brand stringongoingHighest-trust corroboration available

The Wikidata one surprises people. You can’t make a Wikipedia article about yourself without getting reverted, and you shouldn’t try. Wikidata is a different thing: a structured item with properties, notability bar much lower than Wikipedia’s, and it feeds directly into the knowledge graphs several systems consult. Company B has a Wikidata item with 14 properties including official website, country, inception, founded by (linked to their founders’ own items) and industry. It took an afternoon.

4. Ship structured identity data and mean it. Organization schema on the homepage, sameAs pointing at every corroborating profile you just cleaned up, Person schema for each author with their own sameAs array, WebSite with publisher referencing the Organization by @id. The @id discipline is the part people skip and it’s the part that turns a pile of markup into a graph.

{
  "@context": "https://schema.org",
  "@type": "Organization",
  "@id": "https://cadence.co.uk/#organization",
  "name": "Cadence",
  "legalName": "Cadence Software Ltd",
  "url": "https://cadence.co.uk/",
  "foundingDate": "2021-03-15",
  "founder": { "@id": "https://cadence.co.uk/about/#jen-marlowe" },
  "sameAs": [
    "https://www.wikidata.org/wiki/Q00000000",
    "https://www.linkedin.com/company/cadence-analytics/",
    "https://www.crunchbase.com/organization/cadence-analytics",
    "https://find-and-update.company-information.service.gov.uk/company/09876543",
    "https://www.g2.com/products/cadence/reviews"
  ]
}

Validate with Schema Markup Validator (validator.schema.org) rather than Google’s Rich Results Test, because the latter only shows you what Google renders and will silently ignore the parts you care about here. Run your homepage through it and check that the Organization, Person and WebSite nodes actually reference each other rather than sitting as three orphans.

How to tell whether any of this landed

Measurement is where entity work gets abandoned, because the feedback loop is slow and the obvious metrics don’t move. Three things I’d actually track.

Prompt-set citation rate. Write 30 to 50 prompts a real buyer would type, including the ones where you shouldn’t win. Run them monthly across ChatGPT, Perplexity, Claude and Google AI Overviews and record whether you’re named, whether the name is correct, and whether the claim about you is accurate. Manually is fine at that volume, about 90 minutes a month. Profound, Peec AI and Scrunch will automate it from roughly £90/month if you’d rather. The metric that matters early isn’t appearance rate, it’s misattribution rate: how often you’re named but described wrong, or confused with another entity. Watching that fall from 40% to under 10% is the leading indicator, and it moves months before citation volume does.

There’s a cruder test I like more. Ask a model directly: “What is Cadence, the subscription billing analytics company? Who founded it, when, and where are they based?” Then ask it how confident it is and where it got that. Do it in a fresh session with no memory. You’ll get one of three answers: a confident correct answer, a confident wrong answer (usually a bigger entity with a similar name), or hedging. Hedging means your surfaces disagree with each other. Confident and wrong means you have a name collision and need to think seriously about always pairing your brand string with a category qualifier.

Finally, brand-string consistency as an actual audit. Grep your own site for name variants, then check the top 20 external surfaces by hand once a quarter. Company A found 11 distinct brand strings in their own repo, including two in the <title> templates of different page types.

If you want the retrieval-side companion to this, how content gets selected and quoted once you are resolvable, that’s covered in Getting Cited by AI Search. The two halves work together: entity signals get you into the candidate set, content signals decide what gets quoted.

The uncomfortable trade

All of the above is about eight to fifteen days of work for most companies, and roughly none of it feels like content marketing. No briefs, no drafts, no publishing calendar. You’ll be editing a Crunchbase profile and rewriting a 600-word about page and chasing your ops lead about the Companies House SIC description. It’s admin. It’s deeply unglamorous, and it’s very hard to put in a monthly report in a way that looks impressive.

But run the counterfactual honestly. Another twelve posts, at maybe two days each with review cycles, is 24 days of your team’s capacity going into a system that can’t reliably tell who published them. The entity work is cheaper and it compounds: every surface you make consistent raises the confidence score on every piece of content you’ve already published and every one you publish next.

Start with the name audit this week. It’s the one that costs an afternoon and unblocks everything else.