AI Content Marketing
§4 Section 4 of 6 2,852 words · 13 min

Getting Cited by AI Search

Your traffic report for Q3 probably looks fine. Sessions flat, maybe down 8%. Conversions holding. Nothing that triggers an alarm. Then you search for the thing your best-performing guide ranks #2 for, in ChatGPT rather than Google, and a competitor you’ve never worried about gets named in the answer while you don’t appear at all.

That gap between ranking and being cited is the whole problem. Generative engine optimisation is the work of closing it: making your content the thing a language model reaches for when it assembles an answer, not just the thing a crawler indexed. The mechanics overlap with SEO in places and diverge sharply in others, and the divergences are where most content teams are currently losing.

This page covers what actually drives citation, how to instrument it so you’re not guessing, and the production changes worth making in a small team with no engineering budget.

What Determines Whether You Get Cited

There are three separate systems at play and conflating them will cost you months.

Retrieval-augmented answers (ChatGPT with search on, Perplexity, Google AI Overviews, Copilot) run a query, pull a handful of documents, and synthesise. Citation here is a two-stage problem: you need to survive the retrieval step, then be quotable enough to survive the synthesis step. Retrieval is largely conventional search — Perplexity leans heavily on its own index plus Bing, ChatGPT search uses Bing, AI Overviews use Google’s index. If you don’t rank in the top 20 for the underlying query, you’re not in the candidate set.

Parametric answers come from training data alone, no retrieval. You get “mentioned” rather than cited, and the lever is volume and consistency of mentions across the web at training time. This is a two-year game and mostly a digital PR problem, not a content-production one.

Grounded assistant answers with tool access sit in between and increasingly dominate: an agent might run three or four searches, fetch five pages, and reconcile them. Contradictions between sources get resolved in favour of whichever source states things most precisely.

Most of your controllable upside is in the first and third. Which means the question isn’t “how do I write for AI” but “what makes a model pick my paragraph over the other four it fetched.”

From testing across a few hundred queries, the pattern that holds up: models cite passages that answer the question in a self-contained way, contain a specific verifiable detail, and require no context from elsewhere in the page. A paragraph beginning “As we discussed above, this approach has several benefits” is unciteable. A paragraph beginning “Perplexity’s free tier caps Pro searches at five per four hours; the $20/month tier raises this to 600 per day” is highly citeable, because it can be lifted whole and it contains numbers that make the answer feel substantiated.

The Retrieval Layer: Getting Into the Candidate Set

Before any of the writing advice matters, check whether you’re being fetched at all.

Three server-log checks, in order of how often they turn up something broken:

Filter your access logs for GPTBot, OAI-SearchBot, ChatGPT-User, PerplexityBot, Perplexity-User, ClaudeBot, Claude-SearchBot, Claude-User, Google-Extended, Bingbot, Amazonbot, and meta-externalagent. These are separate agents with separate purposes and blocking the wrong one is a common own goal. GPTBot is training-data collection. OAI-SearchBot builds the search index that ChatGPT search queries. ChatGPT-User is the live fetch when a user’s question triggers a page retrieval. A robots.txt that blocks GPTBot on privacy grounds but leaves OAI-SearchBot allowed is a defensible position. One that blocks all three because someone pasted in a blocklist from a Reddit thread in 2024 means you’ve opted out of citation entirely, and nobody on the marketing team knows.

Second: check status codes for those agents specifically. Cloudflare’s bot-fight mode, Akamai’s bot manager, and most WAF default rulesets will happily serve 403s to PerplexityBot while your Googlebot numbers look perfect. If you’re on Cloudflare, the AI Crawl Control panel (formerly AI Audit) under the Security section gives you per-crawler request counts and lets you allow specific agents without touching robots.txt.

Third: fetch your own page the way a live retrieval agent does. ChatGPT-User and Perplexity-User do not execute JavaScript reliably. Run curl -A "Mozilla/5.0 (compatible; ChatGPT-User/1.0; +https://openai.com/bot)" https://yoursite.com/your-best-guide/ | wc -c and then look at what you actually got. If your framework ships a shell and hydrates content client-side, the model sees the shell. I’ve seen a 4,000-word guide return 11KB of nav and footer, which is exactly as useless as it sounds.

There’s a machine-readable surface worth building alongside this, which is covered properly in llms.txt and Machine-Readable Content for Marketing Sites — that page goes into the file format, what belongs in it, and the honest picture of which engines currently read it.

Structuring Pages for Extraction

Once you’re being fetched, the question becomes which chunk of your page makes it into the model’s context and whether that chunk stands alone.

Retrieval systems chunk documents. The exact sizes vary and aren’t published, but typical RAG implementations use 200–800 token windows with some overlap. Your page isn’t evaluated as a page. It’s evaluated as maybe fifteen fragments, and the fragment that gets retrieved is whichever one is semantically closest to the query. This has a direct production consequence: every H2 section needs to make sense to someone who has read nothing else.

What that looks like in practice, with a before and after from a SaaS pricing guide:

Before:

How It Compares

As mentioned, the per-seat model has drawbacks at scale. The alternative approach avoids these by shifting the cost basis, though it introduces the complexity we touched on earlier.

After:

Per-Seat vs Usage-Based Pricing for Support Desk Tools

Per-seat pricing (Zendesk Suite Growth, £69/agent/month) becomes expensive for teams with many occasional users: a 40-person team where 12 answer tickets daily and 28 answer them weekly pays £2,760/month regardless. Usage-based tools like Intercom charge per resolution (from $0.99), which suits the same team at roughly £600/month at 600 resolutions, but costs more than per-seat above ~2,800 resolutions.

The second version is longer and it does more work. It names products, carries numbers, states the crossover point, and answers the question without the reader needing the preceding 1,200 words. It’s also just better for humans, which is the thing that keeps this discipline from becoming a cargo cult.

Practical rules that hold up in testing:

  • Front-load the answer under each H2. The first 40–60 words after the heading should contain the claim. Models do pick up content further down, but the opening sentences carry disproportionate weight in embedding similarity against a question-shaped query.
  • Use question-shaped headings where the query genuinely is a question, but don’t force it. “How much does generative engine optimisation cost?” is fine. “What Is Content?” is not a query anyone types.
  • Keep tables simple. Two-level headers, merged cells, and nested markup get mangled in extraction. A flat table with a header row survives; a table with colspan across three sub-headers usually doesn’t.
  • Put the definitive statement in prose, not only in a graphic. If your comparison chart holds the key number and the surrounding text just says “see the chart,” the number doesn’t exist as far as retrieval is concerned.

On schema: Article, FAQPage, HowTo, Product and Organization markup are worth maintaining, but treat them as retrieval support rather than a citation lever. The evidence that structured data directly increases LLM citation is thin. The evidence that it helps you rank in the conventional index that feeds retrieval is solid, and that’s reason enough. Where schema does earn its keep specifically for GEO is author and sameAs on your bylines, because entity disambiguation is genuinely how models decide whether “Sarah Chen, content strategist” is a person with a track record or a name on a page.

Measurement: Building a Citation Panel

You cannot manage this from GA4. Referral traffic from AI surfaces exists but it substantially undercounts influence, because a large share of AI answers resolve the user’s question without a click. Measure citations directly.

Build a prompt panel. Not 500 prompts; 40 to 60 is plenty for a small team and is actually maintainable. Structure it in four buckets:

  1. Category-definitional (8–10 prompts): “what is generative engine optimisation”, “how does GEO differ from SEO”
  2. Commercial comparison (15–20): “best GEO tracking tools for small teams”, “Profound vs Peec AI”, “alternatives to [your product]”
  3. Problem-shaped (15–20): “my content ranks but doesn’t get cited in ChatGPT, why”, “how do I know if AI crawlers can read my site”
  4. Branded (5–8): “is [your brand] any good”, “what does [your brand] do”

Run each prompt monthly across ChatGPT, Perplexity, Google AI Overviews, Claude, and Gemini. Record four fields: cited yes/no, mentioned-without-link yes/no, which competitors appeared, and which of your URLs was used. That’s 60 prompts × 5 engines = 300 data points, which is four to five hours by hand, or automated.

For automation, the tools have consolidated into a reasonable field. Profound is the enterprise option and prices accordingly (typically $500+/month, often much more). Peec AI sits around €90–120/month and is built for exactly this small-team use case. Otterly.ai starts near $29/month for a limited prompt count. Ahrefs Brand Radar and Semrush’s AI toolkit are worth checking if you already pay for those platforms, since the marginal cost is often zero to low. If you’re technical enough to write 60 lines of Python, hitting the OpenAI and Perplexity APIs directly with web search enabled and logging results to a Google Sheet costs about $15/month in API calls and gives you exactly the fields you want.

What to watch in the numbers: citation rate by bucket, not overall. A healthy pattern for a mid-sized content programme after six months of deliberate work is something like 70–90% on branded, 25–40% on problem-shaped, 15–30% on commercial comparison, and 5–15% on category-definitional. If your branded rate is under 60%, you have an entity problem, not a content problem, and the fix is Wikipedia-adjacent: consistent naming, a Crunchbase entry, LinkedIn company page, G2 listing, and getting described the same way in third-party coverage.

Also log server-side. Set up a GA4 or Matomo segment for referrers matching chatgpt.com, perplexity.ai, copilot.microsoft.com, gemini.google.com, and claude.ai. The absolute numbers will be small and that’s expected. What matters is the conversion rate, which is consistently and substantially higher than organic search across everything I’ve seen reported — roughly 3–6x organic in most accounts, because the user arrives already convinced by the summary and clicks to verify or buy.

Using AI in Production Without Producing Citation-Proof Content

The irony sitting in the middle of this discipline: the fastest route to content no model will cite is generating it with a model and shipping it.

Generic AI-generated content fails at citation for a structural reason. A model asked to write about GEO produces the median of what’s been written about GEO. Retrieval then has no reason to prefer your median restatement over the forty other median restatements, and synthesis has nothing distinctive to lift. You’ve produced a document that is maximally similar to the corpus and therefore maximally redundant.

Where AI earns its place in the workflow:

Gap analysis against actual answers. Take an AI answer to one of your panel prompts, paste it alongside your page on the same topic, and prompt: “List every factual claim in answer A that is absent from document B. For each, state whether B contradicts it, omits it, or covers it less specifically.” This produces a concrete edit list in about 90 seconds. Run it on your top 20 pages and you’ll typically find three to eight fixable gaps per page.

Extraction testing. Paste your page into a fresh context and ask: “Using only this document, answer: [target query]. Quote the exact passage you used.” If the model quotes your intro rather than your section on the topic, your section isn’t self-contained. If it says the document doesn’t cover it, your heading is lying about what’s underneath it.

Interview transcription and structuring. This is the highest-value use and the most underused. The differentiated material in your organisation is in the heads of your solutions engineers, your support leads, and your customers. Thirty minutes recorded with a support lead, transcribed with Whisper or Otter, then structured into a draft, gives you specifics no competitor has. The model does the structuring work; the humans supply the thing worth citing.

Schema and metadata generation. Tedious, rule-bound, low-risk. Automate it fully.

Where AI should not be in the workflow: writing the claim-bearing sentences. The numbers, the comparisons, the “in our testing” statements. These are the citation surface, and a hallucinated figure that gets picked up and repeated by an answer engine is a worse outcome than no citation at all.

A rough allocation that works for a three-person team producing eight pieces a month: AI handles roughly 40% of total hours (research synthesis, outlines, transcription, first-pass structure, metadata, repurposing), humans handle the 60% that includes every sentence carrying a number or a judgement.

What to Do in Your First 60 Days

Days 1–5. Pull server logs for the eleven user agents listed earlier, going back 90 days. Check robots.txt against what you actually intend. Curl-test your top ten pages with a non-JS user agent and confirm the body text is in the HTML. Fix whatever’s broken here first; nothing downstream matters if you’re serving 403s to OAI-SearchBot.

Days 6–15. Build the 40–60 prompt panel. Run it manually once across five engines to get a baseline. This is boring and there’s no way around it. Record competitors carefully, because the single most useful output of the baseline is a list of who’s getting cited instead of you, which tells you what good looks like in your category.

Days 16–40. Rewrite your top 15 pages for extractability. Self-contained H2 sections, front-loaded answers, specific figures in prose. Pull the numbers from your own data where you can: your own benchmark, your own survey, your own support ticket analysis. Original data is the single strongest citation driver because it’s the one thing a model can’t get from the other four sources it fetched.

Days 41–60. Re-run the panel. Expect movement on branded and problem-shaped prompts first, since those have the shortest retrieval-to-citation path. Category-definitional prompts move on a six-to-twelve-month timescale and sometimes never, because they’re often answered parametrically.

The Entity Problem Nobody Wants to Own

Here’s where small teams get stuck, and it isn’t a writing problem.

If a model has no coherent representation of who you are, it won’t cite you even when your page is retrieved, because citing an unknown source in a confident answer is exactly the behaviour these systems are tuned against. Entity strength is doing quiet work underneath every citation decision.

Concretely, for a UK B2B company: a Wikipedia article if you genuinely meet notability (most won’t, don’t waste months on this), a complete and accurate Crunchbase profile, Companies House data that matches your site, a G2 or Capterra presence with real reviews, consistent founder bios across LinkedIn and your about page, and being described in the same terms by third parties. Contradictory descriptions are worse than sparse ones. If your site says “AI content operations platform,” your LinkedIn says “marketing automation,” and TechCrunch called you “a content workflow startup,” a model reconciling those three has low confidence in all of them.

The mechanism to influence here is unglamorous: get mentioned in the roundup posts, comparison pages, and “best X for Y” listicles that answer engines actually retrieve when someone asks a commercial query. When Perplexity answers “best content marketing tools for small UK agencies,” it typically fetches three to six listicles and synthesises. Being in those listicles is the citation. Not your page ranking for that term.

Two hours a month of outreach to the people who maintain those listicles will move your commercial-comparison citation rate more than another 3,000-word guide. It’s the least fashionable recommendation on this page and it’s the one I’d argue hardest for.

When Citation Isn’t the Right Goal

Not every page should chase this. A bottom-funnel case study, a pricing page, a product comparison built to convert: these earn their keep through people who arrive with intent, and optimising them for extraction can dilute what makes them convert.

The pages worth putting through the full treatment are the ones answering questions people ask before they know vendors exist. Definitional content, how-to content, benchmark and data content, comparison content covering your category rather than just your product. That’s maybe 30–40% of a typical content library, and concentrating effort there beats a shallow pass across everything.

Something worth watching as you work through this: the queries where you get cited and the queries where you get traffic are drifting apart. A page can be cited in forty answers a month and send you eleven visits. That’s not failure. It’s a different kind of presence, and the teams who learn to value it correctly will make better decisions than the ones still grading every page on sessions.

In this section

The supporting pages under this subject.