AI Content Marketing
028 Getting Cited by AI Search 1,338 words · 6 min

Schema Markup That Earns You AI Answers

Let’s start with what schema markup for AI search won’t do. It won’t make ChatGPT, Perplexity or Google’s AI Overviews cite you. Nobody outside those companies can honestly promise that, and anyone selling “schema for AI rankings” as a lever is selling a story.

What schema does is narrower and more useful. It removes ambiguity. A model, or the retrieval system feeding it, has to work out what a page is (a how-to? an opinion piece? a product listing?), who wrote it, when, and what it claims. Structured data states those things in a machine-readable form instead of leaving them to inference. If you’ve already done the harder work of being worth citing, as covered in our pillar on getting cited by AI search, schema is the cheap insurance that your page is read correctly.

Here’s the short version of my position: implement four types well, validate them, and ignore the rest.

Why ambiguity is the real problem

Take a UK B2B content team publishing “How to Run a Content Audit”. The page might be a tutorial, a template download, a case study or a sales page. The author byline says “Sam”. Which Sam? The publish date says March 2023, but the page was substantially rewritten last month. The company name appears as “Acme”, “Acme Ltd” and “Acme Digital” in three places.

A human resolves most of that in seconds. A retrieval pipeline working across millions of pages resolves it statistically, and sometimes wrongly. Schema gives it a declared answer. Google’s documentation is explicit that structured data helps it understand page content and entities, and that’s the part you can rely on. Whether another engine uses it the same way is something I can’t verify, so I won’t pretend otherwise.

Use JSON-LD, placed in the page head or body. It’s Google’s recommended format, and it keeps the markup separate from your HTML, which means your developer (or you, in a CMS plugin) doesn’t have to touch templates line by line.

The four types worth implementing

1. Article (or BlogPosting)

This is the foundation. It declares the headline, dates, author and publisher in one block.

{
  "@context": "https://schema.org",
  "@type": "BlogPosting",
  "headline": "How to Run a Content Audit in a Week",
  "datePublished": "2025-03-12",
  "dateModified": "2026-09-04",
  "author": {
    "@type": "Person",
    "@id": "https://www.example.co.uk/team/sam-okafor/#person"
  },
  "publisher": {
    "@id": "https://www.example.co.uk/#organization"
  },
  "mainEntityOfPage": "https://www.example.co.uk/blog/content-audit/"
}

Two habits matter here. First, keep dateModified honest: change it when the substance changes, not when you fix a comma. Second, use @id references rather than repeating author and publisher details inline, so every page points at the same entity.

2. Person (for authors)

The author line is where most content sites are vaguest. Build one Person entity per writer, on a proper author page, and point every article at it.

{
  "@context": "https://schema.org",
  "@type": "Person",
  "@id": "https://www.example.co.uk/team/sam-okafor/#person",
  "name": "Sam Okafor",
  "jobTitle": "Head of Content",
  "worksFor": { "@id": "https://www.example.co.uk/#organization" },
  "sameAs": [
    "https://www.linkedin.com/in/example-sam-okafor",
    "https://muckrack.com/example-sam-okafor"
  ]
}

The sameAs array does the disambiguation. It tells a system which Sam Okafor you mean, by linking to profiles that already exist elsewhere. Two or three solid links beat ten weak ones.

3. Organization

One block, sitewide, usually on the homepage and referenced by @id everywhere else. Include your legal name, logo URL, and sameAs links to your Companies House entry, LinkedIn page and any Wikidata item you have. For a UK business, the Companies House link is a cheap, authoritative anchor that most competitors skip.

4. FAQPage, used narrowly

This one comes with a caveat. In August 2023 Google restricted FAQ rich results to well-known government and health sites, so for most of us the visual payoff in classic search has gone. I still implement it, but only where a page genuinely contains question-and-answer content that’s visible on the page. The markup describes the claims the page makes, in the same words a reader sees. Don’t bolt it onto pages where the questions were invented for the markup. If the visible text and the schema disagree, you’ve created ambiguity rather than removing it.

(HowTo is the cousin people ask about. Google deprecated its rich results in 2023 too. Skip it unless you have a specific reason.)

What wastes developer time

Developer hours are scarce in a small team. Here’s where I’d refuse to spend them.

  • Review and AggregateRating on your own content. Self-serving reviews of your own services are against Google’s guidelines for most cases, and invite a manual action.
  • Speakable. Still marked beta, limited to news publishers in the US, and not something AI engines are documented to read.
  • Deep nesting for its own sake. A twelve-level graph with about, mentions, knowsAbout and isPartOf on every page is maintenance debt. Each property is another thing that can drift out of sync with the copy.
  • Hand-coding per page. If you’re pasting JSON-LD into individual posts, you’ll have stale dates inside six months. Generate it from CMS fields.
  • Anything for an llms.txt-style promise. Different topic, same trap: no major engine has confirmed it reads the file.

Implementation without an engineer

For most teams of one to five, the sensible route is a plugin, then a check.

PlatformToolWhat it covers
WordPressYoast SEO or Rank MathArticle, Person, Organization via a connected graph
WordPress (finer control)Schema ProCustom mappings from fields
WebflowCustom code embed plus CMS fieldsArticle and Person templates
Anything elseGoogle Tag ManagerInjects JSON-LD, but test it carefully

Yoast’s graph output already links Article, Person and Organization with @id, which is exactly the structure above. Your job is mostly filling in author profiles properly, since an empty bio field produces thin markup.

Validate before you trust it

Run every template type through two free tools:

  1. Google’s Rich Results Test (search.google.com/test/rich-results) shows whether Google can parse the markup and flags errors.
  2. Schema Markup Validator (validator.schema.org) checks against the full schema.org vocabulary, including types Google ignores.

Then check Search Console’s Enhancements reports a week or two after rollout. Errors show up there at scale, per template.

A worked example: a 40-post audit

Say you run a UK SaaS blog with 40 posts and a three-person team. A realistic audit looks like this:

CheckTypical findingFix
Author entities6 writers, 2 have no author pageBuild 2 pages, add sameAs
dateModifiedMatches datePublished on all 40Map to the CMS “last updated” field
OrganizationName differs on 3 templatesSingle @id, one legal name
FAQPageOn 11 posts, 4 with questions not visibleRemove from those 4
Errors in Search Console14 warnings, missing imageSet a default featured image

That’s about a day of work for someone comfortable in a CMS, and perhaps two hours of developer time for the template mapping. No heroics.

Measuring it honestly

You can’t isolate schema’s effect on AI citations, because too much else moves at once. Be straight with your boss about that. What you can measure:

  • Error and warning counts in Search Console (target: zero errors).
  • Rich result impressions for the types that still qualify.
  • Whether your brand and author names are being represented correctly when you prompt ChatGPT, Perplexity and Gemini with ten to twenty queries your buyers use. Log the answers monthly in a spreadsheet: cited or not, name correct or not, date correct or not.

If the names and dates in those answers start matching your markup, that’s a real signal that ambiguity went down. If you never get cited, the problem is the content, not the JSON.

Start with the Organization block this week, then authors, then Article. Do FAQPage last, and only where the page earns it.