AI Content Marketing
022 Getting Cited by AI Search 1,840 words · 8 min

Original Data as Citation Bait

Ask Claude or ChatGPT or Perplexity “what’s a good email open rate for B2B SaaS in the UK?” and watch what happens. You’ll get a number, a range, and one or two citations. The citations are almost never a blog post explaining what open rates are. They point to whoever actually counted: a benchmark report, a survey, a platform’s aggregated data.

That’s the whole game now. Synthesis costs nothing. A model can write “five ways to improve your email open rates” better than most agencies, instantly, for free, and it doesn’t need your blog to do it. What it cannot do is invent the fact that 312 UK B2B marketers told you their median open rate was 24.1% in August 2026. That number exists in exactly one place. If a model wants it, it has to name you.

So the question for a small content team stops being “how do we write better than AI” and becomes “what do we know that nothing else knows.” Original research content marketing is the answer, and it’s more achievable on a two-person team than almost anyone believes.

Why proprietary numbers survive and explainers don’t

Think about what’s happening mechanically when an LLM answers a question. It’s either recalling something from training or retrieving live sources and summarising them. In both paths, a factual claim that needs attribution creates pressure to cite. A generic recommendation doesn’t.

Here’s the difference in practice. A model summarising ten posts about content strategy produces an answer with no citations, because every post said roughly the same thing and none of it is a checkable claim. A model answering “how long does it take to rank a new page in 2026” reaches for someone’s dataset, because the honest answer is a number and numbers have owners.

Every explainer you publish is competing against a system that can generate the same explanation on demand. Your 1,800 words on “what is topic clustering” have a replacement cost of about four seconds. Your survey of 400 practitioners has a replacement cost of six weeks and £3,000, which is exactly why it keeps getting cited.

There’s a secondary effect worth understanding: original data gets picked up by other publishers, and those pickups become training data and retrieval sources themselves. You publish a number, three trade publications quote it, eleven blog posts quote those publications. Now the claim exists in dozens of documents across the web, all of them attributing back to you. That’s the compounding mechanism explainer content never gets. We go deeper on the retrieval side of this in Getting Cited by AI Search, but the sourcing logic stands on its own.

The quarterly cadence: what two people can actually ship

Forget the 5,000-respondent industry report. That’s an enterprise play with a research agency attached. What a two-person team can do is one modest, tightly-scoped study per quarter, published properly, on a budget that survives a finance review.

Here’s a real shape for it:

PhaseDaysWhoCost
Question design and screener2Content lead£0
Survey build (Typeform / Tally)1Content lead£0–25/mo
Recruitment (Prolific, 300 responses)3Either£900–1,400
Cleaning and analysis3Analyst-ish person£0
Write-up, charts, methodology4Content lead£0
Outreach and pickup chasingOngoingEither£0

That’s roughly 13 working days spread over five weeks and £1,000–1,500 in hard costs for a study you can cite for two years. Run four a year and you’ve spent under £6,000 building the only asset in your content programme that a competitor cannot copy by writing faster.

Prolific’s pricing works out around £2.50–4.00 per completed response for a five-minute survey with a professional screener, plus their service fee (currently 33% on top of participant payment for the standard tier). Screening hard for “works in marketing at a UK company with 50+ employees” pushes the effective cost up, because you’re paying for screen-outs too. Budget 1.4x your naive estimate.

Worked example one: the survey

Say you sell a content operations tool. Your instinct is to survey people about content operations. Resist it, because “how do you manage your content calendar” produces answers nobody needs to cite.

Instead, find a question where a number is genuinely missing from the public record and where that number is load-bearing in other people’s arguments. Something like: how much of a UK in-house content team’s output is now AI-assisted, and has that changed what they measure?

Your questionnaire, and keep it to nine questions:

  1. Team size (1, 2–5, 6–15, 16+)
  2. Sector
  3. Roughly what percentage of published content involved an LLM at any stage? (0 / 1–25 / 26–50 / 51–75 / 76–100)
  4. Which stages? (multi-select: ideation, outlining, drafting, editing, repurposing, measurement)
  5. Has your published volume changed in the last 12 months? (down / flat / up to 2x / more than 2x)
  6. Has your headcount changed?
  7. Do you disclose AI involvement to readers?
  8. What’s your primary success metric now, and what was it 12 months ago?
  9. Open text: one thing that got harder.

Question 5 crossed with question 6 is your headline. If 41% of teams doubled output on flat headcount, that’s a statistic with legs, and it’s the kind of thing that gets quoted in every “AI and content jobs” argument for the next eighteen months. Question 8 gives you a second story about measurement drift. Question 9 gives you pull quotes, which is what makes the write-up readable rather than a table dump.

One methodological point that separates credible from junk: report your n for every single cut. If you slice by team size and the 16+ bucket has 22 respondents, say so, and say the margin is wide. Every serious publisher who considers quoting you will look for this, and its absence is the fastest way to get ignored by the people whose pickups you need.

Worked example two: the dataset you already have

Surveys aren’t the only route, and often aren’t the cheapest. If you run any kind of platform, tool, agency, or even a moderately busy set of client accounts, you’re sitting on data nobody else has.

An agency with 40 client sites in Google Search Console has a dataset. Pull it through the Search Console API into BigQuery (the bulk export is free up to BigQuery’s usage tier, and 40 sites of moderate traffic sits comfortably inside it), and you can answer questions the public benchmarks can’t:

-- Median days from first impression to first page-1 ranking, by content type
SELECT
  content_type,
  COUNT(DISTINCT url) AS urls,
  APPROX_QUANTILES(days_to_top10, 100)[OFFSET(50)] AS median_days
FROM `agency.gsc_first_rank`
WHERE first_impression_date >= '2025-01-01'
GROUP BY content_type
ORDER BY median_days;

+-----------------+-------+-------------+
| content_type    | urls  | median_days |
+-----------------+-------+-------------+
| product page    |   418 |          61 |
| how-to guide    |   932 |          94 |
| comparison page |   276 |         112 |
| original research|   88 |          38 |
+-----------------+-------+-------------+

That last row is a headline whether or not it’s flattering, and the fact that you can’t control what the data says is precisely what makes it worth citing. Anonymise by aggregating (never name a client, never publish a cut with fewer than about 30 URLs behind it), get client sign-off on the aggregate, and you have a study that cost you three days of SQL and nothing else.

Other datasets hiding in plain sight: your own email platform’s send data across a client base, job postings you’ve scraped from LinkedIn for a role category, pricing pages you’ve tracked quarterly across 60 competitors in a category, support ticket categories over two years. The tool doesn’t matter much. Tally or Typeform for collection, Prolific or Attest for recruitment, Google Sheets or Python with pandas for analysis, Datawrapper for charts because it produces accessible, embeddable output with the numbers still readable as text.

Publishing it so a model can actually use it

This is where most original research gets wasted. A beautiful PDF behind an email gate is invisible to retrieval. Every number that matters must exist as HTML text on a crawlable page.

Concretely:

  • Put the findings in a real HTML table. Not an image of a table, not a chart with the numbers only in the SVG. A <table> with <th> headers that says what the numbers mean.
  • Write a sentence per key finding that stands alone. “41% of UK in-house content teams doubled published output in the 12 months to August 2026 without adding headcount (n=312).” That sentence is quotable out of context, which is the only way it gets quoted at all.
  • Give the methodology its own section with a heading. Sample size, recruitment source, fielding dates, screener criteria, how you handled non-responses. This is credibility infrastructure and it also happens to be exactly what a model looks for when deciding whether a claim is sourced.
  • Date it visibly and version it. “Fielded 4–18 August 2026.” Undated statistics decay into unusable statistics.
  • No gate on the data itself. Gate a deeper cut, gate the raw CSV, gate a webinar. The headline numbers go out free or they don’t travel.
  • Add Dataset structured data where it genuinely applies, and Article schema with a clear datePublished everywhere else.

Then do the unglamorous part. Email the twelve journalists and newsletter writers who cover your space, one at a time, with the single most surprising number in the subject line. Not a press release. One number, one sentence of context, one link. Pickups are what turn a study into a widely-attested fact rather than a page on your site.

The part that’s actually hard

Choosing the question. Everything else on this list is logistics, and logistics are solvable with a calendar and £1,200. The question is where studies live or die.

A good research question has three properties. Someone is currently arguing about it without data. The answer is a number or a distribution, not an opinion. And you have privileged access to the people or records that hold the answer. Miss the third and you’re doing a survey anyone could do; miss the first and you’ll publish something accurate that nobody needs.

Keep a running list. Every time a client asks you something you can’t answer with a citation, write it down. Every time you read a post that says “many marketers report…” with no source, write that down. Within a month you’ll have thirty candidate questions, and the quarterly cadence becomes a matter of picking the best one rather than inventing a topic under deadline.

Your first study will be rougher than you’d like. Field it anyway. The second one is faster, the third one has a returning panel you can re-survey for year-on-year comparison, and year-on-year comparison is where original research stops being a campaign and becomes a franchise.