AI Content Marketing
§4.1 Getting Cited by AI Search 1,762 words · 8 min

llms.txt and Machine-Readable Content for Marketing Sites

Somebody on your team has probably forwarded you a link about llms.txt with a note saying “should we do this?” The honest answer is yes, but it will take you ninety minutes and it is not the thing that will get you cited. The machine-readable work that actually changes whether an AI assistant can quote your page correctly sits in three other places, and almost nobody does it properly.

This page covers all four layers: llms.txt itself, the markdown-endpoint pattern underneath it, structured data, and crawler access. If you want the wider picture of how retrieval and citation work before you get into file formats, the parent guide on getting cited by AI search sets that up.

What llms.txt actually is

llms.txt is a plain markdown file you publish at the root of your domain, so yourdomain.co.uk/llms.txt. It was proposed by Jeremy Howard of Answer.AI in September 2024 to solve a specific problem: developer documentation sites are full of navigation, sidebars and JavaScript, and when a model has a limited context window you want to hand it a clean curated index instead of raw HTML.

The spec is deliberately thin. An H1 with the site name, a blockquote summary, then H2-grouped lists of links in the format [Page title](url): one-line description. An optional ## Optional section marks pages that can be skipped when context is tight. That’s it. No XML, no validation, no submission process.

Here’s the important framing for a marketing team: llms.txt is a hint file, not a directive like robots.txt. Nothing is obliged to read it. Google’s John Mueller has publicly likened it to the keywords meta tag, and as of writing no major model provider has confirmed that their retrieval systems fetch it as a ranking or grounding input. What does fetch it, reliably, is agentic traffic: a user pointing Claude or ChatGPT at your domain, a Cursor or Windsurf session pulling your docs, an MCP-connected research tool crawling for context.

The realistic case for publishing one anyway

Ninety minutes of work for a non-zero chance of being read cleanly is a good trade. The stronger argument is that writing llms.txt forces you to answer a question you should have answered already: if a machine could only read twelve pages of your site, which twelve, and what does each one claim?

That exercise surfaces problems. Most sites I’ve looked at have two pages competing for the same claim, a pricing page that doesn’t state a price, and a “how it works” page that is mostly a video. You cannot write a one-line honest description of a page that doesn’t say anything, so the file acts as an audit.

A worked llms.txt for a small B2B site

This is for a fictional UK SaaS company selling rota software to hospitality groups. Fifty-odd pages on the site, eight of which matter.

# Shiftly

> Shiftly is rota and timesheet software for UK hospitality
> groups with 3-50 sites. It handles split shifts, tronc
> distribution and Working Time Regulations compliance.
> Pricing starts at £4 per employee per month.

## Product
- [How Shiftly works](https://shiftly.co.uk/product): Rota
  building, shift swaps and payroll export in one flow.
- [Pricing](https://shiftly.co.uk/pricing): £4/employee/month
  on Standard, £7 on Pro. No setup fee, monthly rolling.
- [Tronc distribution](https://shiftly.co.uk/tronc): How
  Shiftly splits tips under the Employment (Allocation of
  Tips) Act 2023.

## Guides
- [Working Time Regulations for hospitality](https://shiftly.co.uk/guides/wtr):
  48-hour week, opt-outs, rest breaks, record-keeping duties.
- [Rota templates for pub groups](https://shiftly.co.uk/guides/pub-rotas):
  Four worked patterns with staffing ratios.

## Company
- [About and team](https://shiftly.co.uk/about): Founded 2019,
  Bristol, 34 staff, 1,100 customer sites.

## Optional
- [Changelog](https://shiftly.co.uk/changelog)
- [Careers](https://shiftly.co.uk/careers)

Notice what the descriptions do. They carry facts, not adjectives. “£4/employee/month on Standard” is extractable; “flexible pricing to suit your business” is not. If a model reads only this file and nothing else, it can still answer a pricing question about you accurately. That is the whole point.

Keep the file to 30 links or fewer. Once you’re past that you’re building a sitemap, and you already have one of those.

llms-full.txt and the token maths that kills it

The companion convention, llms-full.txt, concatenates the actual body content of every listed page into one file. For a 40-page API reference this is genuinely useful. For a marketing site it usually isn’t, and the arithmetic explains why.

Take a content programme with 60 published articles averaging 1,400 words. That’s 84,000 words, roughly 112,000 tokens once you add headings and markdown syntax. You have just produced a file that consumes most of a 128k context window and all of the retrieval budget for a task that needed one paragraph. Worse, the file goes stale the moment you update an article, and nothing tells the consumer which version they have.

Publish llms-full.txt if your site is under about 15,000 words total, or if you’re documenting an API. Otherwise skip it and do the next thing instead, which is better anyway.

Markdown endpoints: the part that earns its keep

Serve every page in clean markdown at a predictable URL. Anthropic’s documentation does this: append .md to any docs URL and you get the raw markdown. Mintlify generates these automatically for every hosted site. Several other docs platforms now do the same.

The reason this beats a static bundle is that it’s on-demand and always current. An agent hits yourdomain.co.uk/guides/wtr.md, gets 1,200 words of clean text with no nav, no cookie banner, no consent modal swallowing the first paint, and no JavaScript to execute. Extraction error goes to near zero.

On WordPress you can do this with a small template that returns text/markdown when the request path ends in .md, converting the post content with a library like league/html-to-markdown. On Next.js it’s a route handler reading the same MDX source your page component uses. Either way, budget a day of developer time and reference the .md URLs directly in your llms.txt so they get discovered.

Content negotiation via the Accept: text/markdown header is the more elegant version and worth doing if your CDN makes it easy, but suffix URLs are more likely to actually be tried by a crawler.

Structured data is the layer that’s definitely consumed

Unlike llms.txt, JSON-LD is unambiguously parsed at scale: by Google for AI Overviews and AI Mode grounding, by Bing for Copilot, and by anything built on a search API. It is the least fashionable and highest-yield item on this list.

Three specific things to fix, in order:

Organization, once, sitewide. Include name, url, logo, description, foundingDate, numberOfEmployees, address with the UK postcode, and critically sameAs pointing at your LinkedIn, Companies House listing and Crunchbase. Add knowsAbout with four to six topic strings. This is how a model resolves “Shiftly” to an entity rather than a word.

Article with a real author. Every post needs author as a Person object with its own url and sameAs to a LinkedIn profile, plus datePublished and dateModified in ISO 8601. A string author name is worth almost nothing. A linked Person with a live profile is the difference between an attributed quote and an unattributed paraphrase.

FAQPage on your genuine FAQ content. Google withdrew FAQ rich results for most sites in August 2023, so there’s no SERP reward. The markup is still a clean question-answer pair in a parseable format, which is exactly the shape a retrieval system wants. Fifteen minutes per page.

Validate with validator.schema.org rather than Google’s Rich Results Test, because the latter only reports on types that produce rich results and will stay silent about markup that’s technically fine but unsupported. Screaming Frog’s Structured Data tab will crawl the whole site and flag missing or invalid types in one pass; on a 500-page site that’s a twenty-minute job on the £199/year licence.

Your robots.txt is now a commercial decision

You cannot talk about machine-readable content without deciding who’s allowed to read it. The user agents worth naming explicitly:

AgentPurposeBlocking cost
GPTBotOpenAI model trainingNo visibility impact
OAI-SearchBotChatGPT search indexRemoves you from ChatGPT search
ChatGPT-UserUser-triggered fetchBreaks link-following in chats
ClaudeBotAnthropic crawlingReduces Claude citation
PerplexityBotPerplexity indexRemoves you from Perplexity
Google-ExtendedGemini training opt-outNo Search or AI Overviews impact
CCBotCommon CrawlWide downstream training effect

The trap is treating training and retrieval as one decision. Blocking GPTBot costs you nothing in citations. Blocking OAI-SearchBot costs you ChatGPT entirely. Plenty of sites have done the second while meaning the first.

Check your logs rather than assuming. If you’re on Cloudflare, the AI Crawl Control dashboard gives you per-bot request counts without touching a log file, and it also tells you whether Cloudflare’s own default blocking is silently sitting in front of your llms.txt work. A site that enabled managed AI bot blocking in 2025 and then published llms.txt in 2026 has built a reading room with the door locked.

What to measure, and what you can’t

Server logs answer the first question honestly: is anything fetching /llms.txt, and which agents. Filter your access log for the path and count by user agent over 30 days. If the answer is four requests, two of them yours, you now know the file’s value is the audit it forced, not the traffic.

The citation question needs a different tool. Profound, Peec AI and Scrunch all track whether your brand appears in answers across ChatGPT, Perplexity, Gemini and Google AI Mode, typically from around £100 to £400 a month at the small-team end. Ahrefs Brand Radar bundles a lighter version into existing subscriptions. Pick one, define fifteen prompts a real buyer would type, and take a baseline before you ship any of this, because otherwise you’ll have no way to separate your structured data work from the model updates that land every six weeks.

The honest expectation for a five-page structured data fix plus markdown endpoints on a 60-article site is a measurable change in extraction accuracy, meaning fewer answers that get your pricing or positioning wrong, within about six weeks. Citation volume moves more slowly and depends far more on whether anyone else on the internet mentions you.

Start with the Organization schema this afternoon. It takes forty minutes, it’s the piece that makes every other signal resolvable to an entity, and unlike llms.txt you know for certain something is reading it.