Structuring Pages So AI Assistants Can Lift Your Answers
There’s a specific test I run on any page I’m about to publish now. I copy a single paragraph out of the middle of it, paste it into a blank document, and read it cold. If I can’t tell what it’s about, who it applies to, or what claim it’s making, the paragraph fails. It might still be beautiful prose in context. But an AI assistant building an answer doesn’t have the context. It has the passage.
That’s the whole game with content structure for AI search. Retrieval systems don’t read your article. They chunk it, embed the chunks, score them against a query, and assemble an answer from the winners. Your carefully built argument gets diced into 200 to 500 token segments before anything decides whether you’re worth citing. A paragraph that opens “This is where most teams go wrong” is dead on arrival, because this refers to something three paragraphs up that didn’t make it into the chunk.
Most advice about this stops at “write clearly” and “use headings”. Useful, vague, not actionable on a Tuesday afternoon with four posts in the queue. So here’s the actual shape, and then a full worked rewrite of a post that had every problem.
Four properties that make a passage liftable
Self-containment. Every passage names its own subject. Not “the platform” but “Shopify Plus”. Not “this approach” but “server-side tagging”. You’re writing for a reader who arrived at paragraph nine by parachute.
Claim-then-evidence order. The assertion comes first, the support comes after. Retrieval scoring rewards passages whose opening sentence matches the query’s semantic shape. If your paragraph builds for four sentences toward a conclusion, the conclusion is sitting in the part of the chunk that’s least likely to drive the match, and may not even be in the same chunk.
Question-shaped headings. “How long does technical SEO migration take?” pulls against a real query. “The Timeline Question” pulls against nothing. Headings become chunk boundaries and often get prepended to the chunk as context, so a heading that states a question gives the whole passage a job.
No anaphora across sections. Anaphora is the grammatical term for pointing backwards: this, that, these, it, the former, as mentioned above. Within a paragraph it’s fine and natural. Across a section boundary it’s a hole in the passage where meaning used to be.
Two of those four are things good editors already push for. The other two feel slightly unnatural the first time, which is why nobody does them consistently.
The worked example: a 1,400-word post, rebuilt
A client of mine (B2B SaaS, workforce scheduling, selling into UK hospitality groups) had a post called “Rota Software: What to Look For”. Decent traffic, 340 organic sessions a month, zero AI citations across a fortnight of manual prompt checks in ChatGPT, Perplexity and Google’s AI Mode. I ran 18 query variations. Nothing.
Here’s the original opening of section three, verbatim except for the brand name:
Getting the Balance Right
This is arguably the trickiest part. You want enough flexibility that managers can respond to a no-show at 6am, but not so much that your labour costs drift. Most of the tools in this category handle it badly. They either lock the schedule down completely or let anyone change anything, and neither works when you’re running twelve sites.
Read that as a standalone chunk. What’s the trickiest part? Which tools? What category? A retrieval system sees a passage about balance and flexibility with no anchor to rota software, hospitality, or multi-site operations. It’s unliftable. It also happens to contain the single best insight in the article.
The rewrite:
How much schedule-editing freedom should rota software give site managers?
Multi-site hospitality operators need rota software that allows site managers to fill same-day gaps without granting open-ended edit rights. The failure mode runs in both directions: locked schedules mean a 6am no-show at one site escalates to head office before anyone can cover it, while unrestricted editing lets labour costs drift with no central visibility. For a group running twelve sites, the workable setting is manager-level authority to swap staff within an approved shift budget, with anything that increases total hours routed to area management for approval.
Same insight. 79 words versus 68, so not much longer. But now it names the buyer (multi-site hospitality operators), the product category (rota software), the actor (site managers), and it leads with the claim rather than winding up to it. The heading is a query someone would actually type.
Nine weeks after restructuring the full post this way, that page was cited in 7 of the 18 tracked prompts. Traffic barely moved: 340 to 370 sessions. The citations were the point.
What the chunking actually looks like
If you want to see why this matters rather than take it on faith, run your own page through a chunker. LangChain’s RecursiveCharacterTextSplitter with a 400-token size and 50-token overlap is close enough to production retrieval behaviour to be diagnostic:
from langchain_text_splitters import RecursiveCharacterTextSplitter
splitter = RecursiveCharacterTextSplitter(
chunk_size=1600, # ~400 tokens
chunk_overlap=200,
separators=["\n## ", "\n### ", "\n\n", "\n", ". ", " "],
)
chunks = splitter.split_text(open("post.md").read())
for i, c in enumerate(chunks):
print(f"--- chunk {i} ({len(c)} chars) ---")
print(c[:180].replace("\n", " "))
On the original rota post, chunk 4 began:
--- chunk 4 (1583 chars) ---
neither works when you're running twelve sites. The other thing to
watch is how it handles holiday accrual, because this is where the
cheaper options tend to fall over. They'll track...
A chunk that opens mid-sentence with “neither works” and continues into “this is where the cheaper options” and “they’ll track”. Three unresolved references in 30 words. Nothing in that block tells a retrieval system what product, what industry, what problem.
After the rewrite, chunk 4 opened with the full H3 and the claim sentence. That’s the difference, and it’s visible in your terminal in about four minutes.
Where writers get nervous, and what to do about it
The objection I hear most: “If every paragraph restates the subject, it reads like a robot wrote it.” Fair, and it’s a real risk. The fix is to vary how you re-anchor rather than dropping the anchor.
Compare three ways of opening a paragraph in a piece about GA4 migration:
| Version | Opening | Self-contained? | Reads naturally? |
|---|---|---|---|
| Original | “This is where it gets messy.” | No | Yes |
| Over-corrected | “GA4 migration for e-commerce sites is where GA4 migration gets messy.” | Yes | No |
| Working | “Event-based tracking is where GA4 migrations get messy for e-commerce.” | Yes | Yes |
The working version does its job in one clause and moves on. You’re not adding a sentence of throat-clearing to every paragraph, you’re choosing a subject noun instead of a pronoun in the first sentence. That’s a word-level edit, not a structural one.
The second worry is that question headings make a page look like an FAQ. They do, a bit. Mitigate it by keeping H2s statement-shaped and making H3s the questions, which is roughly what I’ve done here. The H2 carries the narrative spine and the H3s carry the retrieval surface. Readers scanning the page see a structure; retrieval systems see a set of question-answer pairs.
Third worry, and this one’s legitimate: anaphora is how English handles flow. Strip it entirely and your prose gets stiff. The rule isn’t no pronouns, it’s no pronouns whose antecedent lives in a different section. Within a paragraph, point backwards all you like. At the top of a new H2 or H3, start clean. I keep this as a literal checklist item: read the first sentence under every heading with everything above it covered.
A 40-minute retrofit pass
For an existing post that’s already ranking but never cited, here’s the sequence I use. Budget 40 minutes for a 1,500-word piece.
- Headings first, 10 minutes. Convert every H3 to a question a buyer would type. Leave H2s alone unless they’re metaphors (“The Iceberg Problem” has to go).
- First-sentence sweep, 15 minutes. Read only the first sentence under each heading. Replace every leading pronoun or demonstrative with the actual noun. Check the sentence states a claim rather than announcing one is coming. “There are three factors to consider here” is an announcement. “Three factors determine rota software cost: site count, contracted-hours complexity, and payroll integration” is a claim.
- Orphan hunt, 10 minutes. Search the document for “above”, “below”, “as mentioned”, “earlier”, “the former”, “the latter”. Each hit either gets rewritten to name the thing or gets cut. In practice about half get cut with no loss.
- Chunk check, 5 minutes. Run the splitter. Read the first 20 words of each chunk. If any chunk opens with a fragment you can’t identify, the paragraph above it needs a subject.
On a batch of 11 posts I ran this on across two clients, average time was 37 minutes per post and the most common single fix was step 2. Roughly 60% of first sentences under headings began with a demonstrative or a pronoun. That one habit was doing most of the damage.
None of this replaces the wider question of whether assistants trust your domain enough to cite you at all, which is a mix of entity clarity, corroboration across sources, and plain old authority. I’ve written that up separately in Getting Cited by AI Search, and structure without that foundation gets you liftable passages nobody lifts.
The thing nobody tells you about doing this at scale
Once you’ve restructured five or six posts, you’ll notice something irritating: your best-performing passages start to look interchangeable across articles. Every one names the buyer, states a claim, gives a number. That’s what extractability costs you. The prose gets more uniform at the paragraph level.
I’ve made peace with it by pushing the variety upward. Structure the passages rigidly, then vary the sections: one post opens with a failure story, another with a table, another with a query log. The reading experience stays distinct even when the individual bricks are cut to the same size.
And if you want a single sentence to hold onto while you’re editing: a passage that can’t survive being copied out of the page won’t survive being copied out of the page, which is exactly what’s happening to your content every time someone asks an assistant a question you already answered.