AI-Assisted Content Strategy and Planning
Most content teams that “use AI” are actually using it in one place: the first draft. That’s the worst place to start. Drafting is where AI output is most visibly generic, where the editing tax is highest, and where the quality ceiling is lowest. Meanwhile the parts of the job that genuinely benefit — clustering three years of search data, pressure-testing a content brief before anyone writes a word, reconciling GA4 and Search Console into something a director will read — get done by hand or not at all.
An AI content strategy worth the name allocates AI deliberately across the workflow. You decide which jobs it does well, which jobs it does badly, and what evidence you’d accept that it’s working. This page covers how to do that for a team of one to five people with a real content programme already running.
Start by auditing where your hours actually go
Before you buy anything or write a single prompt, spend a fortnight logging your team’s time in fifteen-minute blocks. Toggl or a Notion table both work. Categorise into: research and planning, briefing, drafting, editing, SEO/technical, publishing/CMS, distribution, reporting, stakeholder management.
A typical three-person in-house team running twelve pieces a month comes out looking something like this:
| Activity | Hours/month | % of capacity |
|---|---|---|
| Research and planning | 42 | 9% |
| Briefing | 28 | 6% |
| Drafting | 96 | 20% |
| Editing and fact-checking | 74 | 15% |
| SEO and on-page | 36 | 7% |
| Publishing and CMS | 52 | 11% |
| Distribution and repurposing | 61 | 13% |
| Reporting | 33 | 7% |
| Stakeholder management | 58 | 12% |
The interesting number isn’t drafting at 20%. It’s that publishing, distribution and reporting together account for 31% of capacity and produce almost no differentiated value. Those are the jobs where AI plus a bit of automation gives you back genuine hours without touching the quality of the thing you publish.
Do this audit properly and your AI roadmap writes itself. You’ll usually find two or three tasks eating 8-12 hours a month each that are highly structured, low-judgement, and repeated in the same form every time. Those go first.
Build a content intelligence layer before you build workflows
The reason AI planning output feels generic is that the model has no idea what your company sells, who buys it, what your last forty articles argued, or which of them worked. Fix the input problem and the output problem mostly solves itself.
Assemble a working corpus. Not a brand guidelines PDF: actual artefacts you’d hand a new senior hire in week one.
- Positioning source of truth. One document, 800-1,500 words, covering what you sell, the three competitors you lose to and why, the two or three claims only you can make, and the claims you’re legally not allowed to make. UK teams in finance, health or legal need that last section explicitly, with the FCA or ASA constraint spelled out.
- Customer language file. Pull 40-60 verbatim quotes from sales call transcripts (Gong, Fathom, or just Otter exports), support tickets, G2 reviews, and win/loss interviews. Tag each with segment and stage. This is the single highest-leverage asset you will build, and it takes about six hours.
- Performance ledger. A CSV of every URL you’ve published in the last 24 months with: publish date, primary keyword, current position, monthly clicks, conversions attributed, word count, format, author, and a one-line “what this argued.” Export the quantitative half from Search Console via the API or Looker Studio; fill the qualitative column manually.
- Pattern library. Five to eight pieces you’d be happy to have judged on, with a short annotation on each explaining what makes it good. “Opens with a client number, not a definition. Uses second person throughout. Every claim has a named source.”
Store these somewhere a model can reach. Claude Projects, a custom GPT’s knowledge base, or a NotebookLM notebook all work fine at this scale. If you’re on Claude, Projects with this corpus attached will change your output quality more than any prompt engineering technique you’ll read about.
One agency lead I know measured this directly: same planning prompt, same model, run once cold and once against a Project containing the four assets above. The cold version produced eleven topic ideas, of which two were usable. The corpus version produced nine, of which seven went into the quarter’s calendar. Same twenty minutes of work.
Use AI for the analysis layer of planning, not the ideas layer
Asking a model for content ideas gets you the statistical average of everyone’s content ideas. Asking it to find structure in data you already own gets you something nobody else has.
Cluster your keyword universe properly. Export everything from Ahrefs or Semrush for your seed terms — realistically 3,000 to 15,000 rows. Pull in Search Console query data too, because that includes long-tail terms you already rank for that no tool will show you. Then run clustering on SERP overlap rather than semantic similarity: two keywords belong to the same page if their top ten results share at least three URLs. Keyword Insights does this for about £0.01 per keyword; so does a Python script with the DataForSEO API for roughly a quarter of that. What you get back is a map of how many pages the SERP actually wants, which is nearly always fewer than the number of keywords you were about to build pages for.
A B2B SaaS team I worked with had a 340-keyword list they’d planned as 340 blog posts over eighteen months. SERP clustering collapsed it to 71 genuine clusters. They built 71 pages in five months instead, and the ones that ranked did so faster because each page had the full semantic weight of its cluster behind it rather than being cannibalised by four near-duplicates.
Run gap analysis against your own archive, not just competitors’. Feed the model your performance ledger and ask which clusters you have zero coverage of, which have three or more competing pages, and which pages have declined more than 40% in clicks year on year. That third list is your refresh queue, and refreshes typically return 2-4x what new pages do per hour invested.
Pressure-test the plan before committing. Give the model your draft quarter’s calendar plus the positioning doc and customer language file, then ask specific adversarial questions. “Which of these twelve topics could a competitor publish word-for-word without changing anything?” “Which assume a level of product knowledge our ICP doesn’t have at this stage?” “Where is the same argument being made twice in different clothes?” You will kill two or three items per quarter this way, and killing a bad piece before it’s briefed saves roughly nine hours.
Give each planned piece a strategic job, and make it measurable
A calendar of topics is not a strategy. A calendar where each row has a declared job, a declared reader, and a declared success condition is.
For each piece, force four fields before it enters the queue:
- Job. Pick one: acquire (net-new organic), convince (moves an in-pipeline buyer), retain/expand (existing customer), or enable (sales team uses it directly). Not two.
- Reader state. Where they are before they read, and what specific belief or knowledge changes by the end. “Believes cheap tools are good enough for under 50 users” → “understands the specific failure mode at 30 users.”
- Proof assets required. The named customer, the internal data cut, the SME who needs to be interviewed. If this field is empty, the piece will be generic no matter who or what writes it.
- Success condition with a date. “Position 8 or better for the head term by day 90” or “cited in 15+ sales conversations in Q1, tracked via Gong keyword alert.”
You can use AI to help populate fields 1 and 2 quickly across a whole quarter. Field 3 is a human job and it’s the one that determines whether the output is worth publishing. That distinction — AI for structure, humans for proprietary substance — is the load-bearing principle of everything on this page.
Once those four fields exist, they feed directly into the brief. The brief is where strategy either survives contact with production or evaporates, and it’s the single highest-leverage document in an AI-assisted workflow: the AI content brief template that prevents generic drafts walks through the exact structure, including the sections that stop models defaulting to listicle-shaped filler.
Set the production model per piece, not per team
Teams that declare “we use AI for first drafts” or “we don’t use AI for writing” are both making the same mistake, which is treating a per-piece decision as a policy.
Four modes, chosen at briefing time:
Mode A — Human-only. Founder POV pieces, original research write-ups, anything trading on a named person’s voice or containing regulated claims. AI touches nothing but the outline sanity-check and a final consistency pass. Roughly 15-20% of output.
Mode B — AI-assisted human draft. The writer writes. AI is used for structural challenge (“what objection have I not addressed?”), for turning interview transcripts into organised quote banks, and for alternative phrasings of paragraphs the writer has already written. Best mode for thought leadership and anything where the argument matters more than the coverage. About 30%.
Mode C — AI draft, heavy human edit. Model drafts against a full brief with corpus attached, human rewrites 40-60% of it and adds all proprietary material. Works for comparison pages, integration pages, glossary terms, and most middle-funnel explainers. Budget 90-120 minutes of editing per 1,500 words and be honest that this is not a small number. Around 40%.
Mode D — Templated generation. Programmatic pages from structured data: location pages, “X vs Y” at scale, product spec pages. Generated from a template with real data fields, spot-checked at 10% sample. Only viable when you genuinely have unique data per page. Maybe 10%, and zero for many teams.
Track editing time per mode for a quarter. Most teams discover Mode C costs more total human hours than they assumed and less than they feared, and that the real saving sits in Mode B, where the writer never faces a blank page but the thinking stays theirs.
Watch for the specific failure modes, because they’re predictable
Homogenisation across the calendar. Ten pieces drafted by the same model against similar briefs converge structurally: same three-part framing, same “it’s worth noting,” same tidy tricolon in the intro. Run a quarterly check where you strip bylines from eight of your pieces and eight competitors’ and see whether anyone on the team can sort them. If they can’t, your differentiation is gone.
Confident wrong numbers. Models produce plausible statistics with plausible attributions that don’t exist. The rule that works: every number in a published piece traces to a URL a human has opened, or to an internal data source named in the doc. No exceptions, no “the model said it was from Gartner.” One UK agency built this into their CMS as a required field per statistic and their fact-check time dropped because the burden moved to draft stage.
Brief decay. Briefs get thinner as teams get faster, and thin briefs are precisely what produce generic output. Watch for briefs dropping below your template’s required fields. If the proof assets field is empty on three consecutive briefs, stop and fix the process.
Reader-value drift. The piece covers the topic completely and helps nobody, because completeness is what models optimise for and usefulness isn’t. The test: could a reader do something differently tomorrow having read it? If the answer is no, the piece is an encyclopedia entry wearing a marketing hat.
Measure the programme, not the prompt
Three metrics tell you whether your AI content strategy is working, and none of them is “hours saved.”
Cost per published piece that hits its success condition. Not cost per piece. Total fully-loaded cost (writer time at their day rate, editor time, tool subscriptions, SME hours) divided by the number of pieces that actually did the job declared in their brief. A team publishing 12 a month at £680 each where 4 hit target is running at £2,040 per successful piece. Publishing 8 at £890 where 5 hit target is £1,424. The second team is winning despite looking slower and more expensive on every vanity metric.
Edit distance as a quality proxy. For Mode C pieces, measure what percentage of the AI draft survives to publication. Google Docs version history or a simple diff gives you this. Consistently above 70% surviving means either your briefs are excellent or your editors have stopped caring, and you should know which. Below 30% means the drafting step is costing you more than it saves and you should move those pieces to Mode B.
Assisted pipeline influence, tracked honestly. GA4 alone won’t tell you this. Combine last-touch conversions from GA4 with self-reported attribution (“how did you hear about us” as a free-text field on your demo form, then categorised) and CRM content-touch data if you have HubSpot or similar. Report all three side by side rather than picking whichever flatters the quarter. The gap between them is itself the interesting finding, and directors respect being shown it.
Set up the reporting once and let AI do the assembly. A monthly routine that pulls Search Console API data, GA4 exports and your CRM content-touch report, then drafts the commentary against the previous three months’ reports, turns a four-hour job into forty minutes of checking and editing. The commentary is where it saves time: the model is genuinely good at spotting “this cluster has gained 340 clicks while the pricing page cluster lost 180” and bad at knowing why, which is the right division of labour.
What to do in the next thirty days
Week one: run the time audit, and build the customer language file. Six hours of pulling verbatims will feel like a detour and will pay back within the month.
Week two: export your full keyword universe and cluster it on SERP overlap. Compare the cluster count to your current content plan. Expect the plan to shrink.
Week three: rewrite your brief template so it carries the four strategic fields, and run the next three briefs through it. Track how long drafting takes against the old ones.
Week four: pick your two worst-performing Mode C pieces from the last quarter, measure edit distance, and decide whether those page types belong in Mode B instead.
The teams getting real leverage out of this aren’t the ones with the cleverest prompts. They’re the ones who did the unglamorous work of building a corpus, then pointed the model at problems that have a right answer rather than at problems that need a point of view.
In this section
The supporting pages under this subject.