AI Content Marketing
024 Governance, Disclosure and UK Compliance 1,903 words · 9 min

UK GDPR and Putting Customer Data Into an LLM

Almost everything written about GDPR and AI marketing tools is aimed at the wrong problem. It worries about frontier model training, about whether OpenAI scraped the open web lawfully, about the constitutional status of a large language model. Fascinating stuff. Entirely irrelevant to your Tuesday.

Your actual exposure sits in three routine acts that a two-person content team performs most weeks: pasting a sales call transcript into a chat window, uploading a CRM export to get segment ideas, and building a reusable knowledge base out of customer stories. Each one is a processing operation under UK GDPR. None of them feel like one. That gap is where teams get into trouble, and it closes with about four hours of work.

The three acts, and what actually bites in each

The actWhat you think you’re sharingWhat you’re actually sharing
Pasting a call transcript“Some quotes about the product”Named individuals, their employer, sometimes their health, redundancy, or financial circumstances
Uploading a CRM export“Job titles and industries”28 columns including email, phone, lifecycle stage, and free-text notes a rep wrote in 2023
Training on customer stories“Our own case studies”A corpus you cannot selectively delete when someone asks you to

Take them in order, because the fixes are different and only one of them is about a toggle in a settings menu.

Act one: the call transcript

Here is the scenario. A 47-minute discovery call, recorded in Fathom or Otter.ai, comes out as roughly 8,000 words. You paste the lot into Claude and ask for three case study angles and a pull quote. Reasonable ask. Good use of the tool, honestly, and the output is usually better than what you’d get from a summary.

Now read what you pasted. The header carries the attendee names and their work email addresses. Around minute 19 the prospect explains that their previous marketing manager went on long-term sick leave, which is why the programme stalled. At minute 34 they name a competitor and describe a contract dispute.

That sick leave reference is health data. Under Article 9 it needs a separate condition on top of your Article 6 lawful basis, and “we wanted a case study” is not one. You didn’t ask for it, the speaker volunteered it, and it is now in a third-party system with whatever retention that system applies.

The fix is not to stop using transcripts. It’s to put a redaction step between the recorder and the model, and to make that step mechanical rather than a matter of remembering. What works in practice is a two-pass approach: run the transcript through a local find-and-replace before it goes anywhere, then prompt for the extraction you want.

Speaker A → PROSPECT_1 (Ops lead, 200-person logistics firm)
Speaker B → AE_1
[names, company, emails, phone numbers removed]
[minutes 18:40–20:10 removed: personal circumstances, not relevant]

Four minutes of work. What survives is the thing you actually wanted, which is how a buyer describes their problem in their own vocabulary. What doesn’t survive is the identifiability, which means you’ve also handled Article 5(1)(c) data minimisation without writing a policy about it.

One more thing on transcripts: your recorder is a processor too. Otter.ai has historically used de-identified audio to improve its models unless an administrator changes the workspace setting, and that setting lives at the account level, not the individual level. If your agency runs on someone’s personal Otter Pro account, nobody has checked it. Go and look now rather than after a client asks.

Act two: the CRM export

This one is the most common and the most fixable. You export from HubSpot to understand what your 4,200 contacts actually have in common. HubSpot’s default contact export runs to 28 or so properties, which is 117,600 cells of customer data landing in your Downloads folder and then in a chat window.

Count the columns you need for the job. For “cluster these by job title and industry so I can see which personas I’m under-serving,” it’s three:

job_title,industry,employee_count
Head of Marketing,SaaS,180
Marketing Executive,Manufacturing,55
Content Manager,Professional Services,320

Three columns, 4,200 rows, zero identifiable individuals. The analysis is identical. The exposure is nil, because this is no longer personal data. Article 5(1)(c) says collect and process what is adequate, relevant and limited to what’s necessary, and the ICO reads “necessary” strictly: if the task works without the field, the field wasn’t necessary.

Purpose limitation deserves a mention here too, because it’s the one marketers trip over without noticing. Those contacts gave you their details to receive a whitepaper or to be contacted about a quote. Using the same records for content analysis is a different purpose, and Article 6(4) asks you to assess whether it’s compatible with the original one. Aggregate, non-identifying analysis of your own customer base almost always is. Feeding individual records into a tool that produces individual-level outputs often isn’t. The line is sharper than it looks: if the output names or targets a specific person, you’ve crossed it.

And if you want the shortcut, here it is. A spreadsheet with no names, no email addresses, no phone numbers, no free-text notes, and no account IDs is not personal data, and UK GDPR does not apply to it. Delete columns before you upload. That single habit removes more risk than any vendor negotiation you will ever have.

Act three: the customer story corpus

The third act is where teams create a problem they can’t easily undo. You’ve got 60 customer interviews and 40 published case studies. You load them into a Claude Project, a custom GPT, or a fine-tuned model, so that everything your team writes sounds like your customers rather than like a language model.

The corpus approach is genuinely good for output quality. The compliance problem is Article 17. When a customer emails to say they’ve left the company and want their quotes and name removed, you have one month to comply, and the answer depends entirely on which of those two architectures you chose.

With a retrieval setup, a Claude Project knowledge base or a custom GPT with uploaded files, erasure is a file deletion. Find the source document, remove it, done. The evidence trail is obvious and you can show it.

Fine-tuning is the opposite. Once a person’s words are baked into model weights, you cannot surgically remove them. Your options are retraining from a cleaned dataset or deleting the model, and both of those are decisions you’ll be making under a statutory deadline you didn’t plan for.

So the rule is simple and worth writing into your process document: retrieval, not fine-tuning, for anything containing customer words. Fine-tune on your own published brand copy if you want a house voice. Keep human beings in files you can delete.

Consent is the other half of this. If your case study release form says “for use in marketing materials,” it does not obviously cover “as training material for an AI system.” Adding one line to that form costs nothing and closes the gap: “We may use this material, including with AI writing tools, to produce future marketing content. You can withdraw this at any time by emailing X.” Roughly 30 of the release forms I’ve seen in the last year say nothing at all about it.

The vendor settings that actually move the needle

Most “is this GDPR compliant?” questions about AI tools resolve to three concrete checks, and you can do all three in an afternoon.

Check one: is your account tier trained on? Consumer tiers and business tiers behave differently at every major vendor, and the difference is not cosmetic. ChatGPT Business and Enterprise exclude your content from training by default; ChatGPT Plus has a toggle under Settings, Data Controls that is on unless someone turned it off. Google’s Gemini in Workspace sits under your Workspace terms; the consumer Gemini app does not. Anthropic’s commercial and API terms differ from its consumer ones, and the consumer terms changed materially in 2025, so check the current version rather than what you remember.

Check two: is there a DPA, and does it cover UK transfers? You need an Article 28 processor agreement with every vendor touching customer data. All the major ones publish one; most require you to accept it explicitly rather than applying it automatically. While you’re there, confirm the transfer mechanism for US processing, either the UK Extension to the EU-US Data Privacy Framework with the vendor named on the certified list, or the UK Addendum to the EU SCCs. A vendor with neither is a vendor you don’t use for personal data.

Check three: how long do they keep it? OpenAI’s API retains inputs for 30 days by default for abuse monitoring, with zero data retention available on request for eligible endpoints. Thirty days is fine for most content work. It is not fine if you’ve pasted health data, and that’s another argument for redacting at source.

Worth knowing: these three checks, plus a short Article 35 screening, are also what an enterprise procurement team will ask you for. Using a new AI tool on customer records hits the ICO’s “innovative technology” criterion, and combined with scale it usually means a DPIA is expected. The template is free from the ICO and takes about two hours to fill in honestly. That document, plus a line in your record of processing activities, is most of what a regulator or a client’s legal team wants to see. If you’re building out the wider picture, including how and when you disclose AI involvement to readers and clients, our governance, disclosure and UK compliance guide covers how these pieces fit together.

The one-line ROPA entry nobody writes

Here’s what a sufficient record looks like for a content team using AI tools. It is not long.

Activity: AI-assisted content production
Data categories: pseudonymised customer interview transcripts; 
  aggregated CRM fields (job title, industry, company size)
Lawful basis: Article 6(1)(f) legitimate interests (LIA dated 12/03/2026)
Special category data: excluded by pre-upload redaction
Processors: OpenAI (ChatGPT Business), Anthropic (Claude Team), 
  Fathom (recording). DPAs in place, UK Addendum executed.
Retention: source files 24 months; vendor retention 30 days (API)
Erasure route: source file deletion from knowledge base, ≤5 working days

Eleven lines. If you can produce that when asked, you are in better shape than the overwhelming majority of UK content teams, including ones at companies with legal departments.

The penalty ceiling under UK GDPR is £17.5 million or 4% of global turnover, which is not the number that should motivate you, because it isn’t landing on a five-person content team. The number that should motivate you is 72 hours, the Article 33 window for reporting a personal data breach. Discover on a Thursday that a freelancer pasted your full customer list into a free chatbot account, and you now have until Sunday to work out what happened, what was in it, and whether to tell the ICO. Every hour you spent beforehand deciding what gets redacted and which accounts are approved is an hour you’re not spending during that weekend.

Start with the transcripts. That’s where the special category data hides, it’s the act your team performs most often, and the fix is a find-and-replace.