Building a Brand Voice Brief an LLM Can Actually Follow
Your tone-of-voice deck says “confident, friendly, human.” You paste that into Claude or ChatGPT, ask for a product launch email, and get back something that opens with “In today’s fast-paced digital landscape” and closes with “we’re excited to share.” Confident, friendly, human. Technically the model complied.
The problem is that adjectives are compression artefacts. “Friendly” is a label a human applies after reading something, derived from a hundred concrete choices: contraction density, question frequency, whether you name the reader, how long the average sentence runs. When you hand a model the label instead of the choices, it reconstructs “friendly” from its training distribution, which is the arithmetic mean of every brand blog post on the internet. That’s where the sludge comes from. Not from the model being generic, but from you asking it to be average.
A brand voice ai prompt that works does the opposite. It gives the model things it can check its own output against: counts, bans, structures, and paired examples showing the same idea done wrong and done right. This post is the conversion process, applied to a real deck.
Why adjectives fail and constraints don’t
Language models are extremely good at pattern completion and extremely bad at self-assessment against abstract criteria. Ask one whether a paragraph is “warm” and you’ll get a confident yes regardless of the paragraph. Ask whether a paragraph contains more than two sentences over 25 words and it will count, mostly correctly, and fix them when told to.
I tested this on a fintech client’s newsletter copy across 40 generations, 20 with their adjective-led brief and 20 with a constraint-led rewrite. The adjective version produced “in today’s” or “in an era of” openers in 11 of 20 drafts. The constraint version, which banned both strings explicitly, produced zero. Average sentence length dropped from 22.4 words to 15.1. The number of drafts needing a full rewrite rather than a light edit fell from 14 to 3.
Nothing about the model changed. Only what it could measure itself against.
Step one: mine the deck for what’s actually underneath
Open your tone-of-voice deck. You’re looking for every adjective, value and pillar, and you’re going to interrogate each one with the same question: what did the writer do on the page that made you reach for that word?
Take a real example. A B2B SaaS deck I worked through in July listed three pillars: “Expert but not academic,” “Direct,” and “Optimistic about the work, realistic about the timeline.” Lovely for a workshop. Useless for a prompt.
Here’s what came out after an hour of pulling apart 15 pieces of their best-performing published copy:
Expert but not academic became: use the industry term once, then the plain-English version for the rest of the piece. Never define a term the reader’s job title implies they know (they’re ops managers, they know what a SKU is). No citations in body copy, link instead. Maximum one statistic per 150 words.
Direct became: the main claim goes in the first or second sentence of every section, never the last. No “it’s worth noting that,” “it’s important to remember,” “arguably,” “to some extent.” Questions as headers are banned. Second person outnumbers first person at least 2:1.
Optimistic but realistic became: every capability claim is paired with a constraint or cost in the same paragraph. No superlatives without a number attached (“fastest” needs a benchmark or it’s cut). Never use “seamless,” “effortless,” “simply,” or “just” in front of a verb describing user action.
Three vague pillars, fourteen testable rules. That’s the conversion.
Step two: build the banned list from your own edits
The single highest-leverage component. Most teams write bans from memory, which gets you the obvious ones (delve, tapestry, leverage as a verb) and misses everything specific to your category.
Better method: take your last 20 AI-assisted drafts and the published versions. Diff them. Every phrase you cut, every construction you rewrote, goes on the list. If you use Google Docs, version history plus a manual pass works fine. If you’ve got the drafts in Notion or a shared drive, drop both versions into a long-context model and ask it to list every deletion and substitution as a table, then sort by frequency.
A UK ecommerce team I ran this with found their top removals weren’t the famous AI tells at all. They were: “customers,” used where the brand says “shoppers” (37 instances), the construction “Whether you’re X or Y, we’ve got you covered” (19), “solutions” as a noun for anything (16), and starting sentences with “Plus,” (14). None of that appears on any generic AI-slop wordlist. All of it was costing them 20 minutes an article.
Your banned list should have three tiers, and the distinction matters because a model treats a flat list as equally weighted:
NEVER USE (hard ban, rewrite required):
- "in today's [anything]" / "in an era of" / "in the world of"
- "delve", "tapestry", "landscape" (figurative), "navigate" (figurative)
- "It's not just X, it's Y" and "It's not about X. It's about Y."
- "we're excited to", "we're thrilled to"
- em-dashes (use commas, colons or full stops)
- rhetorical question as a section opener
AVOID (allowed max once per 800 words):
- "seamless", "robust", "powerful", "unlock"
- three-item lists in a single sentence
- sentences opening with a participial phrase ("Building on this, ...")
PREFER INSTEAD:
- "shoppers" not "customers" or "users"
- "set up" not "onboard"
- "costs £X" not "pricing starts from just £X"
The “max once per 800 words” framing is doing real work. Absolute bans on common words make output stilted; a budget lets the model use the word where it’s genuinely the right one and forces it to find alternatives elsewhere.
Step three: numeric structure rules
Models are unreliable at hitting exact word counts but quite good at ranges and ratios. Give them ranges.
From the same SaaS brief, the structure block that made the biggest difference:
SENTENCE AND PARAGRAPH RULES
- Average sentence length: 14-18 words across the piece.
- At least 3 sentences under 8 words per 500 words. Use them
after a long sentence, not in a row.
- No sentence over 32 words. If a sentence exceeds this, split it.
- Paragraphs: 1-4 sentences. At least two single-sentence
paragraphs per 800 words.
- No two consecutive paragraphs may begin with the same part of
speech pattern (e.g. two openers starting with a gerund, or two
starting "The [noun]").
- Contractions on: use "you'll", "it's", "we've" as default.
Full forms only for emphasis.
- Passive voice under 8% of sentences.
That last one, the consecutive-openers rule, is worth stealing on its own. Uniform paragraph openings are the strongest unconscious tell of machine-generated text, stronger than any single word. Human writers working at speed vary their run-ups because they’re thinking about the next idea, not the shape of the sentence. A model optimising for local fluency does not.
Step four: annotated before/after pairs
This is where most briefs stop short, and it’s the part that carries the most signal per token. One paired example teaches more than a paragraph of rules, because it shows the rule applied rather than stated.
Three to five pairs is the sweet spot. Under three and the model over-fits to a single pattern. Over six and you’re eating context for diminishing returns, plus long briefs start getting partially ignored in the middle (the lost-in-the-middle effect is real and it will quietly drop your rules 12 through 20).
Format each one with the annotation attached. The annotation is not optional. Without it, the model copies surface features of the “after” text instead of learning the transformation.
PAIR 1 — opening a feature announcement
BEFORE:
"In today's fast-moving retail environment, inventory accuracy
has never been more critical. That's why we're excited to
announce our new stock sync feature, designed to seamlessly
integrate with your existing workflow."
AFTER:
"Stock sync is live. It pulls counts from Shopify, Linnworks
and your warehouse system every 15 minutes, so the number on
your product page matches the number on your shelf.
Setup takes about ten minutes."
WHY: The before spends 24 words establishing context the
reader already has. The after leads with the fact, names the
actual integrations, gives a real interval and a real setup
time. "Seamlessly" removed: replaced by the mechanism that
makes it seamless.
PAIR 2 — explaining a limitation
BEFORE:
"While our reporting suite is incredibly powerful, some users
may find that certain advanced customisations require
additional configuration."
AFTER:
"Custom report builders are on the roadmap for Q1. Right now
you get 14 preset reports and CSV export. If you need
something bespoke before then, the API will do it, but you'll
need someone who can write a bit of Python."
WHY: Hedging words ("some users may find", "certain") replaced
with the specific gap and the specific workaround, including
its cost. This is the "optimistic about the work, realistic
about the timeline" pillar as actual sentences.
Pull your pairs from real published work, not invented copy. Your best-performing landing page rewritten backwards into a bad version takes ten minutes and teaches the model your actual house style rather than a plausible imitation of it.
Assembling and deploying it
Order matters. Put the banned list and structure rules near the top and the examples at the bottom, with a one-line restatement of the top three hard bans after the examples. Models attend most reliably to the start and end of a long instruction block.
The full brief for that SaaS client runs about 1,100 words. It lives in three places: a Claude Project’s custom instructions, a ChatGPT custom GPT, and a plain markdown file in their Notion that a junior writer pastes into anything else. Keeping it under roughly 1,200 words is deliberate. Past that, compliance on the later rules noticeably degrades, and you’re better off splitting into a core brief plus a format-specific supplement for email or social.
Then run the check as a separate call. Don’t ask the model to write and self-audit in one go, because it will grade its own homework generously. A second prompt does the job: “Check this draft against the brief below. Output a table: rule, pass/fail, the offending text, a suggested fix. Do not rewrite the draft.” That separation caught 6 violations per 1,000 words on first drafts in the fintech test, against 1-2 when the same model was asked to check its own output inline.
Worth saying that the brief is a living document. Every time you edit a draft, ask whether the edit was already covered by a rule that got ignored (fix the rule’s prominence) or reveals a rule you never wrote (add it). Fifteen minutes a fortnight keeps it accurate. The teams who get the most out of this treat the brief as the single artefact that their whole drafting, brand voice and editing workflow hangs off, not as a one-time setup task.
One caution on measurement: don’t judge the brief by whether drafts “sound right.” Judge it by edit time per 1,000 words, tracked before and after. It’s the only number that survives contact with a team that has four other things to ship this week.