AI Content Marketing
§5.1 Measuring AI Content Performance and ROI 2,034 words · 9 min

Building an AI Content Performance Dashboard in GA4 and Looker Studio

Most teams trying to measure AI content end up in the same place: a GA4 property that knows about pages but has no idea which of those pages were drafted by a model, and a spreadsheet of production costs that nobody has updated since March. The dashboard you actually want answers one question on one screen. Does content produced with AI assistance earn its keep compared to content produced without it, per hour of your team’s time?

That is buildable in GA4 and Looker Studio in about a day of setup, plus a discipline habit that costs thirty seconds per published post. Here is the full build.

Decide what “AI content” means before you tag anything

This is the step people skip, and it’s the one that wrecks the data six months in. “AI-assisted” is not a binary. A post where Claude wrote the first draft and an editor rewrote 70% of it is a genuinely different product from a post where a subject-matter expert wrote everything and used a model to tighten headings.

Pick three or four values, write them down, and freeze them. A taxonomy that works for most small teams:

  • human — outlined, drafted and edited by a person. Model use limited to research or grammar.
  • ai_assisted — model produced section drafts from a human brief and outline, human rewrote and added original material (data, quotes, screenshots, opinion).
  • ai_drafted — model produced the full draft from a brief, human edited for accuracy and tone only.
  • ai_repurposed — derived from an existing asset (webinar transcript, report, long post) by a model, human checked.

Four values is the ceiling. GA4 reports start collapsing high-cardinality dimension values into an (other) row once a property crosses its daily row limits, and you do not want your comparison cohorts silently merging. Four is safe forever.

Get the production method into GA4 as a custom dimension

GA4 can’t infer this, so you have to send it. The mechanism is an event parameter on page_view, promoted to a registered custom dimension.

In your CMS. In WordPress, add a custom field to the post editor. Advanced Custom Fields does this in two minutes with a select field named production_method on the Post type, or you can register a simple meta box if you’d rather not add a plugin. In Webflow, add a plain-text field to the Blog Posts collection. In Ghost, use an internal tag like #ai-assisted and parse it.

In your theme or GTM. Output the value into the data layer before the GTM container snippet fires:

<script>
  window.dataLayer = window.dataLayer || [];
  window.dataLayer.push({
    'production_method': 'ai_assisted',
    'publish_date': '2026-04-12'
  });
</script>

Create a Data Layer Variable in GTM named dlv - production_method. Then open your Google tag (the GA4 configuration tag) and add it under Configuration settings as a parameter named production_method with the variable as its value. Setting it there means every event on the page inherits it, not just page_view, which matters when you start looking at scroll depth and outbound clicks.

In GA4. Go to Admin → Data display → Custom definitions → Create custom dimension. Name it “Production method”, scope Event, parameter production_method. You have 50 event-scoped slots on a standard property, so this is not a scarce resource.

One critical property: custom dimensions are not retroactive. GA4 will show (not set) for every page view collected before you registered it. There is no backfill, no historical patch, nothing. Register it today even if the rest of the dashboard waits a month.

For pages published before you started tagging, you can either leave them as (not set) and exclude them, or backfill the CMS field and accept that only their post-tagging traffic is classified. The second option is better and takes an afternoon for a few hundred posts.

Fix the two GA4 settings that will otherwise ruin the dashboard

Data retention. Admin → Data settings → Data retention. The default on a standard property is 2 months. Change it to 14 months. This governs how far back Explorations and any non-standard dimension combination can reach, and content performance is a 6-to-12-month question. Like the custom dimension, this is not retroactive.

Reporting identity. Admin → Reporting identity. If Google signals is on and you leave the identity on Blended, GA4 applies data thresholding: rows with small user counts get withheld entirely so individuals can’t be identified. On a content dashboard segmented four ways by production method, you will see whole cohorts vanish from a monthly view. Switching to Device-based removes thresholding at the cost of cross-device stitching. For content measurement, that trade is worth it.

If you’re still working out which outcomes the dashboard should be accountable for in the first place, the wider framework in Measuring AI Content Performance and ROI covers the attribution and ROI-model side; this page assumes you’ve settled on your metrics and want them on a screen you’ll actually open.

The cost side lives in a Google Sheet

GA4 knows nothing about what a post cost you. Keep a single Sheet, one row per published URL, with these columns:

url_pathtitlepublish_dateproduction_methodbrief_hoursdraft_hoursedit_hoursasset_cost_gbp

Log hours honestly, including the brief. Teams consistently under-count briefing time on AI-assisted work, which is exactly where the time goes. A good brief for a model-drafted post takes longer than a good brief for a freelancer.

Add a calculated column, total_hours, and a blended internal rate. If your content lead costs £48,000 plus 20% overhead, that’s roughly £30 an hour at 1,900 working hours. Use one rate for everyone rather than trying to be precise; you’re comparing cohorts, not doing management accounting.

Store url_path as the bare path, lowercase, with a leading slash and no trailing slash: /blog/ai-content-briefs. Consistency here saves you a blend debugging session later.

Blending the Sheet with GA4 in Looker Studio

In Looker Studio, add two data sources: the Google Analytics connector pointed at your property, and the Sheets connector pointed at your cost sheet. Then Resource → Manage blends → Add a blend.

The join is a left outer join with GA4 on the left, joined on page path. The friction is that GA4’s path may not match your sheet exactly. Use the Page path and screen class dimension rather than Page path + query string, which drags UTM parameters in and shatters your join. Then create a calculated field in the GA4 source:

LOWER(REGEXP_REPLACE(Page path and screen class, '/$', ''))

Join that field to url_path. Blends take up to five tables, so you have room to add Search Console later.

For the Search Console connector, choose the URL Impression table (not Site Impression) so you get a landing page dimension to join on. Remember it holds 16 months maximum and its clicks will never match GA4’s sessions, because they measure different things and de-duplicate differently. Use it for impressions and CTR, not for traffic.

Two calculated fields carry most of the analysis:

Days live:   DATE_DIFF(CURRENT_DATE(), publish_date)
Age cohort:  CASE
               WHEN Days live <= 30 THEN '0-30'
               WHEN Days live <= 90 THEN '31-90'
               WHEN Days live <= 180 THEN '91-180'
               ELSE '181+'
             END

Without age normalisation the dashboard just tells you that older posts have more traffic. Every comparison table should be filtered to one age cohort, usually 91-180, which is where organic content has typically found its level.

The four pages that earn their place

1. Cohort scorecard. Four scorecards across the top, one per production method, each showing organic entrances per post at 91-180 days. Underneath, a bar chart of engagement rate and one of key event rate by method. Filter control for date range and one for topic cluster.

2. Efficiency. A table with production method as the dimension and calculated metrics: entrances per production hour, key events per production hour, cost per key event. This is the page that changes decisions.

3. Page-level detail. Every URL, with production method, days live, entrances, engagement rate, average engagement time, key events, total hours, and cost per key event. Sortable. Add conditional formatting so anything with engagement rate under 40% turns amber.

4. Quality flags. Three small tables built from filters, not new data. Pages with over 500 Search Console impressions and CTR under 1.5% (title problem). Pages with over 200 entrances and average engagement time under 45 seconds (the classic signature of thin model output that ranks then fails). Pages past 120 days live with zero key events.

A worked read: what 62 posts showed

Here is the pattern from a B2B SaaS team running this setup across 62 posts, measured at 91-180 days live, using median rather than mean because two outliers would otherwise carry everything:

PostsMedian entrancesEngagement rateAvg engagement timeKey event rateMedian hoursKey events per hour
human1824061%2m 14s0.9%11.50.19
ai_assisted3121058%1m 58s0.8%4.20.40
ai_drafted136444%0m 51s0.2%1.10.12

Read the last column. Human posts perform best per post and ai_assisted posts perform nearly as well at roughly a third of the time, producing more than twice the conversion value per hour invested. Fully model-drafted posts look cheapest until you notice they generate the fewest key events per hour of anything on the list, because 1.1 hours of work producing 0.128 key events is worse economics than 11.5 hours producing 2.16.

That result is not universal. It is what this team’s data said, and it justified a specific decision: stop shipping ai_drafted posts entirely, move that capacity into ai_assisted, and keep human for the six or seven pieces a year that carry original research.

Note the one metric worth watching for: Looker Studio’s GA4 connector has no MEDIAN function. To get medians you either export to BigQuery (the sandbox is free, standard export covers up to 1 million events a day) or pull a monthly extract into Sheets and compute there. Averages on content traffic are close to meaningless once a single post takes off.

Where the dashboard will lie to you

Consent mode is the big one for UK properties. If your CMP denies analytics_storage for 30% of visitors, GA4 is modelling or missing that share. Behavioural modelling only kicks in when the property has at least 1,000 daily events with consent denied for seven continuous days, plus 1,000 daily users with consent granted on seven of the previous 28 days, which most small-team content sites never reach. The saving grace is that this bias applies equally across your cohorts, so relative comparisons hold even when absolute numbers don’t.

Selection bias is the subtler one. Teams naturally hand harder, more strategic topics to humans and easier topics to the model. If human posts win, check whether they won because of the process or because they were allocated the better keywords. A quick test: filter the page-level table to a single topic cluster where both methods were used, and see whether the gap survives.

What to do in the first fortnight

Register the custom dimension and change data retention to 14 months today, because both start their clocks from the moment you save them. Spend week one backfilling the production_method field for everything published in the last 12 months, and build the Sheet with honest hours for the last 20 posts rather than guessing at 200. Build pages one and three of the dashboard first; efficiency and quality flags can wait until you have enough tagged posts to populate them.

Then leave it alone for 90 days. The temptation to read the dashboard weekly is strong and the data will not mean anything yet, because organic content hasn’t finished doing what it’s going to do. Put a recurring calendar entry at the 90-day mark, and use the interval to tag properly rather than to check.