Choosing Your Stack: One Model, Two Tools, No Sprawl
Count the AI subscriptions on your company card. If you run a content team of one to five people and the number is above three, you are almost certainly paying for overlap, and worse, you are paying for shallowness. The typical small-team ai content tool stack in 2026 looks like this: ChatGPT Plus for drafting, Jasper for “brand voice”, Surfer or Clearscope for briefs, Descript for repurposing video, a Notion AI seat because it came bundled, and something with a Chrome extension that nobody can quite remember signing up for. Six tools, roughly £280 a month, and output that still reads like it was written by a committee of nobody.
The argument here is narrow and I’ll commit to it: you will get better content from one frontier model used deeply, with at most two supporting tools, than from six specialist wrappers used at surface level. Not because the wrappers are bad software. Because depth in one model compounds and breadth across six does not.
What “used deeply” actually means
Shallow use is a prompt box. You type a request, you get 600 words, you rewrite most of it, you conclude AI is overrated for anything above listicle grade.
Deep use is different in kind. It means the model holds your actual context: the last eighteen months of published posts, your positioning doc, the transcript of the customer call from Tuesday, the search console export showing which queries you already rank for and which you’ve abandoned. Claude’s current context window sits at one million tokens on Opus 5. That’s roughly 750,000 words, or about 1,500 typical blog posts. Gemini 2.5 Pro is in the same territory. You can put your entire content archive in front of the model and ask it what you have never said.
Here’s a worked example from a B2B SaaS team I’d describe as typical: three marketers, £14k monthly content budget, 140 published posts. They exported every post to markdown, dropped the lot into a single project, and asked one question: which of our claims appear in more than five posts, and which of those have we never supported with data? The answer came back in ninety seconds and listed eleven repeated assertions, four of which were doing real work in sales conversations with nothing behind them. That’s a research brief no keyword tool generates, because no keyword tool has read your corpus.
A specialist SaaS wrapper cannot do this. Not because its underlying model is weaker (it is often the same model, marked up), but because the wrapper’s product design constrains what you can put in. Jasper’s brand voice feature ingests a sample of your writing and extracts a style descriptor. Useful. But it’s a compression of your corpus into a few hundred tokens of instruction, not access to the corpus itself. You are buying a lossy summary of your own material and paying £39 a seat per month for the privilege.
The three tests that matter
Feature lists are worthless for this decision. Every tool ships the same features within a quarter of each other. Test on three axes instead.
Context handling. Not the advertised window, the usable one. Ask a concrete question: can I put 200,000 words of my own material in and get an answer that references specific passages accurately? Run the test yourself with something you know cold. Load a long document, then ask about a detail buried three-quarters of the way through. Frontier models handle this now; most wrappers truncate aggressively to control their own API costs and never tell you they’ve done it. A wrapper that silently trims your 40-page positioning document to the first 4,000 tokens will produce confident, plausible, wrong output, and you will not know.
Data terms. Read the DPA, specifically two clauses: whether your inputs train the provider’s models by default, and where processing happens. Anthropic and OpenAI both exclude business-tier API and enterprise inputs from training by default. Many wrappers sit on top of those APIs but add their own retention layer, and their terms are frequently vaguer. For UK teams, check whether the provider offers EU or UK data residency, whether there’s a UK International Data Transfer Addendum in place, and how long prompts are retained (30 days is common; some wrappers say nothing at all). If a tool’s privacy page doesn’t mention sub-processors, you have a procurement problem, not a tooling one.
Exit cost. The one everybody skips. If you cancel tomorrow, what leaves with you? Prompts written in a wrapper’s proprietary template system do not port. Custom GPTs and Claude Projects are more portable than they look, because the instructions are plain text you can copy, but a brand-voice model fine-tuned inside a vendor’s platform is gone the day you stop paying. Score this bluntly: if the answer to “can I export this as text” is no, the tool is renting you leverage rather than building it.
Here’s the test as a scoring sheet worth actually filling in:
TOOL: ____________________ Monthly cost: £______
Context handling
Usable window (tested, not advertised): ______ words
Silent truncation observed? Y / N
Cites specific passages correctly? Y / N
Data terms
Inputs excluded from training by default? Y / N
UK/EU residency available? Y / N
Retention period stated in writing? ______ days
Sub-processors listed? Y / N
Exit cost
Prompts/instructions exportable as text? Y / N
Outputs exportable in bulk? Y / N
Anything here I can't rebuild elsewhere? ____________
Verdict: keep / consolidate / cancel
Three “no” answers in the data terms block and you have something a finance director will ask about eventually. Better it’s you asking now.
So what does the stack look like
One frontier model, paid tier, used as the centre of gravity. Claude or ChatGPT, and the choice between them matters far less than the decision to pick one and go deep. Budget £17 to £25 per seat monthly at UK pricing, or move to API access if your volume justifies it.
Two supporting tools, chosen because they do something the model genuinely cannot:
A data source the model has no access to. Ahrefs, Semrush, or Google Search Console (free, and underused). The model cannot see search volume or your impression data. It can analyse that data brilliantly once you paste it in. A Search Console export of queries where you rank positions 8 to 20, handed to the model alongside the actual post, produces a better optimisation brief than any AI SEO tool I’ve tested, because it combines your real performance data with full-text understanding of what you wrote.
A distribution or production tool with a real technical moat. Descript for video and podcast editing, because transcription-linked editing is genuine engineering. Or your CMS integration. Not a “content optimiser” that scores your draft against a competitor average, which is a feature, not a product.
That’s three line items. Call it £120 a month for a three-person team against the £280 of scattered subscriptions, but the money is the small part. The real gain is that everybody on the team is building fluency in the same system.
The compounding argument
Six tools means six mental models, six prompt idioms, six sets of quirks. Nobody on a three-person team gets good at any of them. You stay permanently at the level where AI output needs total rewriting, which is exactly the level at which people conclude it doesn’t work.
One tool used by three people for six months is different. Someone discovers that giving the model three examples of your best-performing posts plus one you actively dislike produces sharper tone matching than any style instruction. Someone else builds a project that holds the ICP research and stops having to re-explain the audience every session. These learnings transfer between teammates because they share a substrate. The team’s collective skill curve bends upward instead of flat-lining across six shallow competencies.
This is also where the automation layer becomes worth building rather than buying, and it’s the natural next step once your stack is consolidated: custom GPTs, Claude Projects, and repurposing workflows are things you assemble from your own accumulated context, and we go into the mechanics of that in Automation, Repurposing and Custom GPTs.
Where this argument has limits
Two honest caveats, because pretending otherwise would make the rest less credible.
If your team publishes video at volume, the production tooling is not optional and not replaceable by a chat interface. Descript, Opus Clip, whatever fits. That’s a real exception.
And if you have a genuine compliance requirement (regulated financial services, healthcare claims), a tool with audit logging and approval workflows earns its cost in a way that a general model does not. Writer.com exists for a reason. Most small teams do not have this requirement and buy as though they do.
Everything else on your subscription list should be defending itself against the scoring sheet above. Run it this week on the tool you use least. My prediction: you cancel it, nothing breaks, and the £39 goes toward getting genuinely good at the one that’s left.