AI Content Marketing
018 Governance, Disclosure and UK Compliance 1,964 words · 9 min

Copyright, Training Data and Who Owns the Draft: The UK Position

A client asked me last spring whether they owned a 2,000-word pillar page. Their freelancer had drafted it in Claude, restructured it by hand, rewritten roughly half, and invoiced £450. The client’s lawyer had sent back a note citing the US Copyright Office’s 2023 Zarya of the Dawn decision and concluded the page was in the public domain. The lawyer was wrong, and wrong in a specific way that matters to anyone running a UK content programme: they were applying American law to a British asset.

Most of what ranks for AI content copyright UK questions is written for an American reader, or by someone who has not read section 9(3) of the Copyright, Designs and Patents Act 1988. The UK is one of a small handful of jurisdictions (Ireland, New Zealand, Hong Kong, India and South Africa are the others) that legislated for machine-made works before anyone thought it would be needed. That provision, plus the near-total absence of a commercial text and data mining exception, puts your team on different ground from a team in Chicago. Worth knowing where the edges are.

Section 9(3) exists, and it is weirder than you want it to be

Here is the whole thing:

In the case of a literary, dramatic, musical or artistic work which is computer-generated, the author shall be taken to be the person by whom the arrangements necessary for the creation of the work are undertaken.

Section 178 defines “computer-generated” as generated by computer “in circumstances such that there is no human author of the work.” So the provision only fires when human input is negligible. If you wrote a 60-word brief, rejected four drafts, rewrote the intro and restructured the H2s, section 9(3) is not your route. You have an ordinary literary work with a human author, protected for life plus 70 years, with moral rights attached. That is the good outcome, and it is the one you want to be able to prove.

Section 9(3) is the fallback, and the fallback is thin. Computer-generated works get 50 years from the end of the calendar year in which the work was made (section 12(7)), not life plus 70. They carry no paternity right (section 79(2)(c)) and no integrity right (section 81(2)), so nobody has to credit you and nobody is restrained from mutilating the work. Fifty years is plenty for a product comparison page. The point is that you are claiming a different, weaker thing.

Then there is the problem of who “the person by whom the arrangements necessary” actually is. The only real authority is Nova Productions v Mazooma Games (2007), a case about a pool video game, where the court held the arrangements were made by the programmer who wrote the game, not the player pressing buttons. Read that across to a content workflow and the argument runs: OpenAI made the arrangements, you pressed a button. No UK court has tested this on a generative model. Do not bet a content library on the outcome.

Training data: no exception, and the case that went nowhere

Section 29A of the CDPA permits text and data analysis, but only for non-commercial research, and only where the researcher already has lawful access. It cannot be contracted out of (section 29A(5)), which is a nice bit of drafting that helps academics and does nothing for you. There is no commercial equivalent. The 2022 IPO proposal to create one was withdrawn in February 2023 after the creative industries objected; the code of practice that replaced it collapsed a year later; the December 2024 consultation on an opt-out model drew north of 11,500 responses and produced no legislation. The Data (Use and Access) Act 2025 required reports rather than reform. As you read this, no broad commercial TDM exception is in force in the UK.

Which sounds ominous for anyone using a commercially trained model. In practice, the liability sits upstream, and the Getty litigation showed how hard it is to move. Getty Images sued Stability AI in the English High Court, went to trial in June 2025, and dropped its primary training and output infringement claims mid-trial because it could not establish where the training had happened. The November 2025 judgment dismissed the secondary infringement claim on the basis that model weights do not store copies of the works. Getty came away with a narrow trade mark finding about watermarks.

Nobody in your position is being sued over training data. What you can actually be sued over is an output that reproduces a substantial part of a specific protected work, which is a normal infringement question with a normal answer, and which mostly bites on images and on long verbatim passages.

What the vendor terms give you (and what they cannot)

Every major vendor now assigns output rights to the customer. OpenAI’s terms assign all right, title and interest in Output. Anthropic’s commercial terms do the same. Midjourney gives you ownership of assets you create as a paid subscriber, with the wrinkle that any company grossing over $1,000,000 a year must be on a Pro or Mega plan for that to hold. Adobe trains Firefly on licensed Adobe Stock and offers IP indemnification to enterprise customers. Microsoft’s Copilot Copyright Commitment and Google Cloud’s generative AI indemnity cover paying customers who leave the built-in guardrails on.

Read those clauses carefully and you will notice they all assign whatever rights exist. None of them create a right. If the output is a computer-generated work under section 178, the vendor is assigning you a 50-year, moral-rights-free interest whose ownership is contestable under Nova. If it is not protected at all, the assignment transfers nothing. The contract is useful for stopping the vendor from claiming the work. It does not make the work yours against the world.

A worked example

Take a real shape of asset: a 1,400-word “best X for Y” comparison page, produced by a two-person team in a week.

ComponentHow it was madeWhat you hold
Keyword cluster and outlineHuman, from Ahrefs and 6 sales callsOrdinary copyright in the outline as a literary work; the facts underneath are free
9 product summaries, 110 words eachModel draft, light human editContestable. Likely computer-generated under s.178
Comparison table, 9 rows × 6 columnsHuman selection and arrangement of factual dataCopyright in the selection and arrangement; possibly database right too
Intro and verdict, 340 wordsHuman-writtenFull ordinary copyright, life plus 70
4 pull quotes from customer interviewsHuman, transcribedCopyright sits with the speaker unless assigned; get the release
Page as a wholeHuman editorial assemblyOrdinary copyright in the compilation

The page is protected. The weak spot is the 990 words of product summaries, which on a bad day you would be arguing about under section 9(3). Fixing that costs about 40 minutes of a writer’s time: rewrite the summaries with a human hand, keep the version history, and the ambiguity disappears. That is a better use of 40 minutes than a policy document.

The evidence trail is the asset

UK ownership turns on demonstrating human authorship, so the operationally useful move is making that demonstrable rather than asserting it. Three habits do most of the work.

Keep drafts in a system with version history. Google Docs revision history, Notion page history, or a Git repo if your team is technical. A document that appears fully formed in one paste at 14:32 looks like a computer-generated work. One that accumulates across 31 revisions over two days does not.

Log the brief. A 200-word brief with your angle, your customer objections, your structure and your banned phrases is itself a protected literary work, and it is the clearest evidence that a human directed the creative choices. Store it with the asset, not in a Slack thread that expires.

Record which tool and which model. “Claude Opus 4.5, 12 October 2025” takes four seconds to type into a custom field and answers the question a lawyer will ask in two years. This is the same discipline that belongs in your broader disclosure and record-keeping policy, covered in more depth in our governance, disclosure and UK compliance guide.

What to put in freelancer contracts

If you commission writing, the default position under UK law is that the freelancer owns the copyright, and you get an implied licence for the purpose you commissioned it for. That is already a problem before AI enters. Add AI and you have a second problem: the assignment clause in your standard agreement probably refers to “works authored by the Contractor”, which is precisely the wrong language for a work the statute says has no human author.

Five things need to be in the contract. An assignment that covers both ordinary copyright and any section 9(3) interest, with the “arrangements” language stated explicitly. A written waiver of moral rights, because section 87 requires the waiver to be in writing and signed. A disclosure obligation with a defined threshold rather than a vague promise. A warranty and indemnity against third-party infringement. And a tool allowlist, which doubles as your confidentiality control, because a freelancer pasting your unreleased campaign brief into a consumer chatbot tier is a data problem long before it is a copyright one.

Something close to this, adapted by your own solicitor:

7.2  Assignment. The Contractor assigns to the Client with full title
     guarantee all present and future copyright and all other
     intellectual property rights in the Deliverables, including
     (a) any copyright in which the Contractor is the author, and
     (b) any copyright in a computer-generated work within the
     meaning of section 178 of the Copyright, Designs and Patents
     Act 1988 in which the Contractor is or may be taken to be the
     person by whom the arrangements necessary for the creation of
     the work were undertaken.

7.3  Moral rights. The Contractor irrevocably and unconditionally
     waives all moral rights under Chapter IV of the Copyright,
     Designs and Patents Act 1988 in the Deliverables.

7.4  Disclosure. Where more than 25% of the words in a Deliverable
     as delivered originate from a generative AI system without
     subsequent substantive human revision, the Contractor shall
     state this in writing on delivery, naming the system and model
     version used.

7.5  Records. The Contractor shall retain drafting records for each
     Deliverable for 24 months and provide them to the Client within
     5 working days of a written request.

7.6  Permitted tools. The Contractor shall use only the generative
     AI systems listed in Schedule 3 and shall not input Client
     Confidential Information into any system on a consumer or free
     tier, or any system whose terms permit training on inputs.

The 25% figure in 7.4 is a number you should argue about internally before you adopt it. Some teams set it at 50%, some require disclosure at any level. What matters is that it is a number rather than an adjective, because “substantial AI assistance” is unenforceable and everyone knows it.

Rates are worth a thought here too. If you are paying £280 a post and the contract now demands disclosure, record retention and an indemnity, expect either a rate conversation or quiet non-compliance. The teams I have watched get this right built the record-keeping into the delivery template so it costs the writer about two minutes per piece, and left the rate alone.

Pull your last ten commissioning agreements this week and search them for the word “author”. Whatever you find will tell you how much of your library you can currently prove you own.