aeotime

Technical

What Is llms.txt? Format, Example, and Whether It Works

What is llms.txt? The format, a full llms.txt example, how to create one, how it differs from robots.txt, and an honest look at who actually reads it.

An AI-readiness audit, check by check · illustrative
Alex HoldAI search research, aeotime
Published
Last updated
Reading time
7 min read

llms.txt is a markdown file you place at /llms.txt on your site to give large language models and AI agents a short, curated map of your most useful content, with links to clean, LLM-friendly versions of those pages. It was proposed by Jeremy Howard on September 3, 2024, and the specification at llmstxt.org was updated to version 2 on August 10, 2026. It is useful for documentation and agent workflows. As of September 2026, though, no major AI search engine documents it as an input for deciding which sites to cite, and Google says you do not need it to appear in AI Overviews.

Below: the format, a complete example, how to create one, how it compares with robots.txt, and what the evidence says about adoption.

What is llms.txt?

Websites are built for browsers: navigation, scripts, cookie banners, layout markup. An LLM with a limited context window spends much of it on that noise. llms.txt addresses a narrow problem: when an AI system (or a developer pointing an AI tool at your site) wants to understand what you offer, give it a short index in plain markdown and point it to the pages that matter.

The llmstxt.org proposal describes it as a curated overview for LLMs that sits alongside, not in place of, robots.txt and sitemap.xml. A sitemap lists every URL for crawlers. robots.txt sets access rules. llms.txt picks the handful of pages worth reading and says in plain language what each one is.

The v2 update states the intended consumption model directly: agents view or search the llms.txt to find what they need, then follow the links, which should point to LLM-friendly content.

What goes in an llms.txt file?

The specification defines a fixed order of elements:

  1. An H1 with the name of the project or site. This is the only required element.
  2. A blockquote with a short summary containing the key information needed to understand the rest of the file.
  3. Zero or more markdown sections (paragraphs, lists) with more detail. No headings here.
  4. Zero or more H2 sections containing file lists. Each list item is a markdown link, optionally followed by a colon and a note about the page.

By convention, a section titled ## Optional holds secondary links an agent can skip when context is short. Version 2 keeps this as a convention but dropped the tooling that gave it a mechanical meaning.

The spec also recommends serving markdown versions of your pages. You can append .md to the page URL (/pricing.html.md) or replace the extension (/pricing.md); v2 accepts both. Version 2 also adds two discovery hints you can put in an HTML <link> element or an HTTP header: rel="alternate" type="text/markdown" pointing to a page's markdown version, and rel="describedby" pointing to the relevant llms.txt.

llms.txt example

Here is a realistic llms.txt for a fictional B2B invoicing product. The structure follows the specification; the company and URLs are made up.

# Ledgerly

> Ledgerly is invoicing and accounts-receivable software for B2B companies with 10–500 employees. It sends invoices, chases late payments automatically, and syncs with QuickBooks, Xero and NetSuite. Plans start at $49/month; there is a 14-day free trial.

Key facts:
- Ledgerly is not a payroll or expense tool.
- Supported currencies: USD, EUR, GBP, CAD, AUD.
- Data is hosted in the US and EU (customer's choice).

## Product

- [Features overview](https://ledgerly.example/features.md): What Ledgerly does, organized by workflow (invoicing, reminders, payments, reporting)
- [Pricing](https://ledgerly.example/pricing.md): Plans, limits and what each tier includes
- [Integrations](https://ledgerly.example/integrations.md): Accounting and CRM integrations and what data syncs

## Docs

- [Quickstart](https://ledgerly.example/docs/quickstart.md): Send a first invoice in 10 minutes
- [API reference](https://ledgerly.example/docs/api.md): REST endpoints, authentication, rate limits
- [Webhooks](https://ledgerly.example/docs/webhooks.md): Event types and payload examples

## Comparisons

- [Ledgerly vs manual invoicing](https://ledgerly.example/compare/manual.md): Time and error comparison for teams sending 100+ invoices a month

## Optional

- [Changelog](https://ledgerly.example/changelog.md): Release notes since 2024
- [Security overview](https://ledgerly.example/security.md): Certifications and data handling

A few choices worth copying:

  • The blockquote carries the facts an AI most often gets wrong: category, who it is for, starting price, what it is not. If a model reads only the first 60 words, it should still describe you correctly.
  • Every link has a note that says what the page answers, not a marketing tagline.
  • Links point to .md versions, so the agent gets clean text instead of rendered HTML.
  • It is short. An llms.txt listing 400 URLs is a sitemap in the wrong format.

For a real-world reference, compare the files AI labs publish for their own developer docs, such as OpenAI's and Anthropic's.

How to create llms.txt

  1. Pick the pages. Choose the 10–40 pages that best explain what you are, what you sell, and how to use it: product overview, pricing, key docs, comparisons, policies. Skip blog archives and tag pages.
  2. Write the H1 and blockquote. Use your exact brand name. In the summary, state category, audience, core use case and one or two hard facts (price, availability).
  3. Add a short facts section. Plain bullets for anything commonly misunderstood: what you do not do, where you operate, supported platforms.
  4. Group links under H2s. Product, Docs, Guides, Comparisons, and Optional for the rest.
  5. Write one-line notes for each link. Describe the question the page answers.
  6. Publish markdown versions of the linked pages if your stack allows it. Static site generators and docs platforms often can; for a CMS, a plugin or a build step that strips layout is enough.
  7. Serve it at /llms.txt as plain text, returning HTTP 200. Add the rel="describedby" link if you want to follow v2 fully.
  8. Validate it, then regenerate it whenever key pages change, ideally as part of your build.

You can draft a file from your sitemap with our free llms.txt generator and check structure and broken links with the llms.txt validator. According to llmstxt.org, several platforms can also generate one for you, including Mintlify, GitBook, Yoast SEO, AIOSEO and Wix.

llms.txt vs robots.txt: what is the difference?

They solve different problems and do not conflict.

robots.txt llms.txt
Purpose Tell crawlers which URLs they may access Give AI systems a curated guide to your best content
Format Plain-text directives (User-agent, Allow, Disallow) Markdown (H1, blockquote, link lists)
Standard RFC 9309, followed by major search and AI crawlers Community proposal at llmstxt.org
Controls access Yes, for crawlers that honor it No
Documented by AI search crawlers Yes (OpenAI, Anthropic, Perplexity, Google all document their robots.txt tokens) No, as of September 2026
Affects Google AI Overviews Yes, blocking Googlebot removes you Google says it is not needed

Google's own robots.txt introduction notes that robots.txt is mainly for managing crawler traffic and is not a mechanism for keeping a page out of Google; use noindex for that. The practical rule: robots.txt decides whether AI crawlers can read your site; llms.txt, at best, helps an agent that is already reading find the right pages faster.

If your goal is visibility in ChatGPT search, for instance, the lever that OpenAI documents is robots.txt: its crawler page says sites that opt out of OAI-SearchBot will not be shown in ChatGPT search answers. Check what your robots.txt allows with the AI crawler checker.

Does llms.txt help SEO?

Not for Google, by Google's own account. The AI features documentation for AI Overviews and AI Mode says you do not need to create new machine-readable files, AI text files or markup to appear in those features. Eligibility is the same as for regular search: indexed and eligible to be shown with a snippet.

Google's John Mueller was blunter in June 2025. Search Engine Roundtable reported his Bluesky post: "FWIW no AI system currently uses llms.txt." He also noted that server logs show AI crawlers fetching pages, not the llms.txt file.

The other engines are quieter rather than negative. OpenAI's crawler documentation, Anthropic's crawler help page and Perplexity's bot documentation all explain their user agents and robots.txt handling, and none of them mentions llms.txt as something their search crawlers read. So "llms.txt SEO" as a ranking or citation tactic has no documented support.

Who actually reads llms.txt?

The honest picture as of September 2026:

  • AI labs publish llms.txt for their own docs. OpenAI, Anthropic and Google's Gemini API docs all have one. That signals the format is useful for developer documentation, not that their search products read yours.
  • Chrome's Lighthouse checks for it. An agentic-browsing audit looks for /llms.txt. A 404 counts as "not applicable" because the file is optional; only server errors get flagged. The page does not mention search ranking.
  • Coding assistants and agents use it on demand. When a developer points an AI tool at a docs site, or an agent browses a site to complete a task, a clean index and markdown pages save tokens and reduce errors. This is where the format demonstrably works.
  • Search crawlers mostly ignore it. No major engine has documented llms.txt as an input to crawling, ranking or citation.

Should you add an llms.txt file?

It depends on what you sell.

  • Worth doing now: developer tools, APIs, SaaS with substantial documentation, anything people will use AI coding assistants or agents with. The file costs an hour and helps real users.
  • Low priority: local businesses, publishers, small ecommerce sites. It will not hurt, but it will not move AI visibility either.
  • Never a substitute for: crawlable HTML, correct robots.txt rules, clear answer-first pages, and a consistent brand description across the web. Those are what AI search engines actually retrieve. See how to optimize content for AI search for the work that moves citations.

Quick checklist

  • robots.txt allows the AI search crawlers you want (fix this first)
  • /llms.txt returns 200 with an H1, a factual blockquote and 10–40 annotated links
  • Linked pages have markdown versions, or at least clean HTML
  • The file is regenerated when pricing, product or docs change
  • Validated with the llms.txt validator
  • Expectations set: helps agents and developers, not rankings

If you want to know whether AI engines already mention you, with or without an llms.txt, run the free AI visibility checker.

Frequently asked questions

Where should the llms.txt file go?

At the root of your site, so it loads at https://yourdomain.com/llms.txt. Version 2 of the proposal also allows files in subpaths, such as /docs/llms.txt, which cover the pages under that path; when several apply, the most specific one wins.

What is llms-full.txt?

llms-full.txt is not part of the llmstxt.org specification. It is a convention some documentation platforms, such as Mintlify, generate alongside llms.txt: a single markdown file with the full text of the docs, so a developer can load everything into an AI tool with one URL.

Can llms.txt block AI crawlers?

No. llms.txt has no access-control function. To allow or block specific AI crawlers, use robots.txt rules for their user agents (for example GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot or Google-Extended), and firewall rules if you need enforcement.

Can an llms.txt file hurt my SEO?

There is no evidence it does. It is a plain text file that Google says you do not need for its AI features. The realistic risk is maintenance: a stale file that links to removed pages or old pricing can mislead an agent that reads it.

Written by

Alex Hold

AI search research, aeotime

Alex writes about AI search visibility at aeotime — how ChatGPT, Perplexity, Gemini, Claude and Google AI Overviews choose which brands to mention, which sources they cite, and what site owners can change.

Keep reading

Find out what AI says about you.

Free check · no signup · or start a 7-day trial