AI Content Marketing
011 Measuring AI Content Performance and ROI 1,457 words · 7 min

The Metrics That Change the Moment AI Enters Production

A content lead at a UK B2B SaaS company showed me her 2026 Q1 board slide last spring. Output up from 6 posts a month to 19. Cost per post down from £420 to £110. Blended organic sessions up 4%, which is inside the noise band for a site that size. Every number on the slide was true and the slide told her nothing she could act on.

That gap is the whole problem with most ai content metrics dashboards right now. They were built when drafting was the expensive step, so they measure drafting. Once a competent writer with Claude or ChatGPT can produce a structurally sound 1,400-word draft in forty minutes, counting drafts is like counting keystrokes. The constraint moved. Your measurement did not move with it.

Cost per post was always a proxy, and now it’s a broken one

Think about what cost per post was actually standing in for. It was a rough measure of how much finite human attention each published asset consumed: research, drafting, editing, sign-off, upload. Drafting used to be 50 to 60% of that, so cost per post tracked total effort reasonably well.

Strip drafting down to near zero and the proxy collapses. The remaining costs (subject-matter interviews, fact-checking, legal review, the editor’s judgement about whether this asset should exist) barely shift, and in an assisted workflow some of them go up. Reviewing a fluent draft that’s wrong in three places takes longer than reviewing a clumsy draft that’s right. Every editor who has worked with AI output for six months knows this. Almost none of them have a number for it.

Volume has the same disease. Publishing 19 posts instead of 6 only matters if the 13 extra posts do something, and “do something” is a downstream question with a three-to-six-month lag. Meanwhile you need weekly signals about whether the production system is healthy.

Three measures do that work. They’re cheap to instrument, they’re leading rather than lagging, and they break in informative ways.

Metric one: per-asset yield

Per-asset yield counts publishable assets produced per unit of original research. Not per draft. Per unit of new input, which for most teams means an interview, a data pull, a product deep-dive or a genuinely original point of view.

Here’s a real shape of it. A two-person content team records a 45-minute customer call, transcribes it in Descript, and runs the transcript through a structured extraction prompt. What comes out over the following week:

AssetCountPublished?
Case study, 1,800 words1Yes
LinkedIn posts (founder + brand)3Yes
Newsletter section, 600 words1Yes
Proof blocks for pricing page2Yes
Help-centre FAQ answers43 of 4
Sales one-pager1Yes
Webinar talking points1No

Yield: 11 published assets from one research unit. Before they instrumented anything, the same call produced a case study and, if someone remembered, a LinkedIn post. Yield 1.4.

Notice what this metric rewards. It pushes you towards doing the expensive, un-automatable thing (talking to a customer) and then extracting everything from it, rather than towards generating twenty posts about topics nobody researched. Teams optimising for volume commission twenty briefs. Teams optimising for yield book four more interviews and mine them properly.

Track it in whatever already holds your calendar. In Airtable, add a research_unit_id field to your content table and a rollup that counts published records per unit. A monthly average of 6 to 8 is strong for a small in-house team; under 3 usually means your assisted output is being generated from thin air rather than from anything you know and competitors don’t.

Metric two: editorial rework rate

Rework rate is the share of an assisted draft that doesn’t survive editing. It’s the single most useful number in an AI-assisted content programme and hardly anyone measures it, because measuring it requires keeping the draft.

So keep the draft. Save the assisted output as a separate file before anyone touches it, then diff it against the published version at a word level. A 30-line Python script using difflib does this; so does git diff --word-diff if your drafts live in a repo. Output looks like this:

$ python rework.py drafts/2026-09/pricing-comparison.md published/pricing-comparison.md

draft words                1,412
published words            1,530
retained verbatim            408   28.9%
lightly edited (>60% sim)    211   14.9%
rewritten or cut             793   56.2%
------------------------------------------
rework rate                        71.1%

Run that across a quarter and the pattern is more useful than any single figure. One agency team I’ve seen numbers from found rework rates of 12% on customer-led case studies (the interview did the work, the model just structured it), 34% on technical how-tos where the SME reviewed the outline first, 58% on product comparison posts, and 71% on category-level thought leadership.

Read that distribution as a commissioning instruction, not a scorecard. At 12%, assisted drafting is genuinely saving hours. At 71%, your editor is paying a fluency tax: rewriting confident prose is slower than writing from a blank page, and the team should draft those pieces human-first. Two of the four categories above got moved out of the assisted workflow entirely, and the team’s throughput went up, because the editor stopped spending Tuesdays rescuing thought-leadership drafts.

A working threshold: above 55%, stop assisting the drafting step for that content type and use AI upstream instead (research synthesis, outline pressure-testing, SERP gap analysis in Ahrefs or Semrush, internal-link suggestions).

Metric three: survival rate of assisted output

Survival rate is the share of assisted drafts that reach publication. Count briefs at the top and published URLs at the bottom, within a fixed window such as 30 days.

Running the SaaS team’s numbers from that board slide: 24 briefs commissioned in March, 19 drafts produced, 11 published inside 30 days. Survival 45.8%. Their human-first baseline from the previous year was 82%.

Those eight abandoned drafts are not free. Each consumed a brief, a generation cycle and (critically) an editorial review before anyone decided to kill it. Fold that back into cost and the picture changes:

nominal cost per published post      £110
+ abandoned draft cost (8 × £48)     £384  → £35 per published post
+ editorial review on kills          4.5 hrs → £158 → £14 per post
-----------------------------------------------------------------
true cost per published post         £159

Still better than £420, and not remotely the 74% saving on the slide. More importantly, the gap between nominal and true cost is the diagnostic. A survival rate under 60% almost always means briefs are too vague for the model to hit, and the fix is upstream: sharper briefs, a named source, a stated point of view, an actual audience decision the piece is meant to influence.

Low survival also tells you something uncomfortable and useful. If nearly half of what you generate isn’t good enough to publish, the bottleneck is editorial capacity, and buying more generation capacity will make things worse. Budget the editor.

What to put in front of the board instead

Keep your outcome metrics exactly as they are. Assisted sessions, conversions by landing page in GA4, query-level impressions in Search Console, pipeline attribution in HubSpot: none of that changes because a model helped write the draft, and the temptation to invent AI-specific outcome measures should be resisted. If you need to connect production health to commercial results, the pillar on measuring AI content performance and ROI covers the attribution side in detail.

What changes is the operational layer underneath. Report per-asset yield, rework rate by content type, and survival rate, each with a prior-quarter comparison, and you’ve given your leadership something with decisions attached to it. Yield falling means the research pipeline is starved. Rework climbing in one category means move that category out of the assisted flow. Survival dropping means briefs or capacity, and you can tell which by looking at where drafts die.

Instrumenting this in a fortnight

Week one is plumbing, and it’s genuinely small. Add three fields to your content tracker (research_unit_id, assisted: y/n, draft_saved_at). Create a folder convention so pre-edit drafts persist, drafts/YYYY-MM/slug.md, and make saving the raw output a step in your Notion or Asana template rather than something people remember. Write the diff script or borrow one.

Week two you get your first honest baseline, and it will probably be worse than you expect. That’s the point: 71% rework on your flagship content type is a number you can do something about on the Monday after you see it, in a way that £110 per post never was.