llms.txt is a Markdown file, usually served at /llms.txt, that gives AI agents a short, annotated list of a site's most useful pages. It is often recommended as a way to show up in AI answers, but the public record is more modest: the format is an informal proposal, Google Search says it ignores the file, and the crawler documentation from OpenAI, Anthropic and Perplexity covers robots.txt, not llms.txt.
That does not make it useless. For documentation sites and APIs it is a cheap index that agents and developer tools can read on demand. This guide covers the format, how it differs from robots.txt and sitemap.xml, when it is worth adding, a minimal example and what matters more for AI answers. Provider statements reflect their public pages as of 1 October 2026.
Scope of this guide
Provider statements reflect their public documentation as of 1 October 2026. Sitelemetry's llms.txt check confirms only that the file is reachable and looks like a link list; it does not validate the format, test the links or show whether any AI system reads it.
What the llms.txt proposal specifies
Jeremy Howard published the proposal on 3 September 2024. The current text at llmstxt.org is version 2, modified on 10 August 2026. It is an informal overview kept in a public GitHub repository and open for community input, not an IETF or W3C standard. No search engine or AI provider is obliged to read it.
The file is Markdown, with its parts in a fixed order:
- An H1 with the site or project name. This is the only required part.
- A blockquote with a short summary of the key facts.
- Optional paragraphs or lists with more context, without headings.
- Zero or more H2 sections with file lists: list items that hold a link, optionally followed by a colon and a note, such as
- [Pricing](https://example.com/pricing): plans and limits.
A section named Optional holds secondary links that an agent can skip when it needs a shorter context.
The file can sit at the root or under a path such as /docs/llms.txt. It covers the URLs under its path, and when several files apply, agents should use the most specific one. The proposal does not use /.well-known/, because well-known locations exist only at the origin root and many authors control only a path on a shared host.
The proposal also suggests clean Markdown copies of pages at the same URL, with .md appended (page.html.md) or, since version 2, with the extension replaced (page.md). Version 2 adds two link relations, sent as HTML <link> elements or an HTTP Link header: rel="alternate" type="text/markdown" for the copy and rel="describedby" for the llms.txt that covers the page. The current text does not define a companion file such as llms-full.txt; when a tool or documentation platform produces one, that is its own convention.
How llms.txt differs from robots.txt and sitemap.xml
The three files do different jobs for different readers.
| File | What it does | What it does not do |
|---|---|---|
| robots.txt | Tells cooperating crawlers which URLs they may request (Robots Exclusion Protocol, RFC 9309). | It is not access control. The RFC says its rules are not a form of access authorization. |
| sitemap.xml | Lists the URLs you consider important so search engines can discover them. | A listed URL is not necessarily crawled or indexed. |
| llms.txt | A short, annotated list of key pages that an agent reads when it needs information. | It does not allow or block crawling, does not affect indexing, and nobody is required to read it. |
The proposal draws the same line: robots.txt says what automated access is acceptable, while llms.txt is used on demand. It adds that a sitemap is no substitute, because it often lacks Markdown versions of pages, does not include URLs on other sites and usually covers more content than fits in a model's context window.
So if you want to allow or block an AI crawler, edit robots.txt. Leaving a page out of llms.txt does not hide it, and listing a page there does not override a robots.txt rule that blocks a crawler. The robots.txt guide explains how rules are matched.
What Google and AI providers have said about llms.txt
Here is what the primary sources say:
- Google Search. Its guide to generative AI features (updated 10 July 2026) says you do not need machine-readable files, AI text files or Markdown to appear in Google Search, including its AI features, because Search does not use them. It names llms.txt directly: keeping one for other services that use such files is fine, but it neither helps nor harms visibility or rankings in Google Search.
- Chrome Lighthouse. The experimental Agentic Browsing category (Chrome 150 or later, no 0–100 score) has an llms.txt audit that checks for the file at the domain root. A server error is flagged; a 404 is marked not applicable, because providing the file is optional. That is a presence check in a developer tool, not a statement that any search or AI product uses the file.
- OpenAI. Its crawler page explains robots.txt controls for OAI-SearchBot (search results in ChatGPT) and GPTBot (training), and notes that user-initiated ChatGPT-User visits may not follow robots.txt. It does not say that OpenAI reads a site's llms.txt.
- Anthropic. Its help article lists ClaudeBot, Claude-User and Claude-SearchBot and says they honor robots.txt. It does not mention llms.txt.
- Perplexity. Its documentation covers PerplexityBot for search results and the user-triggered Perplexity-User fetcher, which generally ignores robots.txt. It does not say that Perplexity reads a site's llms.txt.
- Microsoft Bing. No Microsoft or Bing guidance on llms.txt was found. The February 2026 announcement of AI Performance in Bing Webmaster Tools does not mention the file.
As of 1 October 2026, none of these providers documents using a site's llms.txt for ranking, citations or answers. Several of them publish one for their own developer documentation (OpenAI, Anthropic, Perplexity and Google's Gemini API docs), as does Cloudflare. Publishing a file for your own docs is not the same as reading other sites' files.
When llms.txt is worth adding, and when it can wait
The file is cheap to write and, according to Google, harmless in Search. The real cost is keeping it correct.
| Type of site | Priority | Why |
|---|---|---|
| Developer documentation, API references, SDKs | Worth adding | Developers point coding agents at docs. A curated index saves the agent from guessing which page to read. |
| Software products with many setup or integration pages | Worth adding | A short map of setup, limits and pricing helps an agent answer how-to questions from current pages. |
| Open-source projects | Worth adding | The docs are often Markdown already, so the build can produce the file and the page copies. |
| Small business and brochure sites | Low priority | A few well-linked pages are already easy to read. |
| Shops, news sites, blogs | Low priority | Content changes constantly, and a hand-maintained list goes stale. |
A quick test: would a developer paste your URL into a coding assistant and ask it to work with your product? If yes, an llms.txt with Markdown copies of the key pages is an afternoon well spent. If you want to be named in answers about a local service, spend that afternoon on clear page content and crawler access instead. If agents should act on your product rather than read about it, look at an API or an MCP server; the MCP introduction explains how that works.
A minimal llms.txt example for a small site
A complete file for a hypothetical scheduling app at example.com:
# Example Rota
> Example Rota is a browser-based app for scheduling shifts in small cafés and shops. A free plan covers one location; paid plans add more locations.
Example Rota does not run payroll. Last reviewed: 2026-10-01.
## Docs
- [Getting started](https://example.com/docs/getting-started.md): Create a location, add staff and publish a first week
- [Shift templates](https://example.com/docs/templates.md): Repeating shifts, breaks and cover rules
- [API reference](https://example.com/docs/api.md): REST endpoints, authentication and rate limits
## Product
- [Pricing](https://example.com/pricing): Plans, limits and what each plan includes
- [Data and privacy](https://example.com/privacy): Where data is stored and how to export or delete it
## Optional
- [Changelog](https://example.com/changelog): Release notes by month- Serve it as text.
https://example.com/llms.txtshould return status 200 and the Markdown itself. Single-page apps sometimes answer unknown paths with their HTML shell, so check the response withcurl. - Markdown copies are optional. If you do not publish them, link the normal HTML pages, as the Product section above does.
- Keep the summary factual. State what the product is, who it is for and its important limits. Leave slogans out.
- Write notes, not prompts. Asking a model to recommend you adds nothing a reader can check.
- Be selective. Ten good links beat two hundred; the full list of URLs belongs in the sitemap.
With Markdown copies, each HTML page can point to its copy and to the file that covers it:
<link rel="alternate" type="text/markdown" href="/docs/getting-started.md">
<link rel="describedby" href="/llms.txt">Generators, checkers and validators: keeping the file accurate
A file that lists an old price or a deleted page can be worse than none, because an agent that follows it may repeat outdated facts.
- Build it with your docs. If the documentation is generated from Markdown, produce llms.txt and the page copies in the same build step.
- Put it on the release checklist. When prices, plan names, platforms or URLs change, update the file in the same release.
- Check the links automatically. Every URL should answer 200 without a chain of redirects. This shell one-liner prints the status code of every absolute link in the file:
curl -s https://example.com/llms.txt | grep -oE 'https://[^) ]+' | while read -r u; do echo "$(curl -s -o /dev/null -w '%{http_code}' "$u") $u"; doneAn llms.txt generator is fine for a first draft built from a sitemap or a docs folder. It cannot decide which pages matter, and it may include pages set to noindex or blocked in robots.txt. Review the output like any page you publish.
An llms.txt checker or validator can confirm the mechanical parts: the file is reachable, the H1 comes first and the links resolve. No checker can tell you whether an assistant reads the file, because the providers do not document that.
What matters more for AI answers
If you want to be a usable source for AI answers, these come first:
- Crawl access for the bots you want. AI crawlers are controlled in robots.txt, each under its own token. OpenAI separates OAI-SearchBot (ChatGPT search) from GPTBot (training), Anthropic separates Claude-SearchBot from ClaudeBot, and Perplexity documents PerplexityBot. The Google-Extended token is a robots.txt control, not a separate crawler: it governs whether content Google crawls may be used for Gemini training and grounding and, according to Google, does not affect inclusion or ranking in Google Search. Allowing search crawlers while blocking training crawlers is a legitimate choice. Keep in mind that a
User-agent: *group withDisallow: /blocks every crawler that follows robots.txt and has no group of its own. OpenAI says a robots.txt change can take about 24 hours to reach its systems. - Authentication for private content. User-triggered fetchers such as ChatGPT-User and Perplexity-User may not follow robots.txt, and RFC 9309 says robots rules are not access authorization.
- Clear content on the page. Google's guidance for its AI features points to unique, useful content and a clear technical structure that is ready for discovery and indexing, not to special files. Answer real questions directly in the HTML of a reachable page. The AI visibility guide works through an example.
- Structured data that matches the page. Google says structured data is not required for its generative AI search and that no special schema.org markup is needed; it still recommends structured data for rich-result eligibility in Search. If you use JSON-LD, keep it consistent with the visible facts.
- Consistent facts. Name, prices, plans, regions and contact details should agree across your pages, your structured data and your llms.txt. If your own pages disagree, any system summarizing them has to guess.
No file can make an assistant cite you. What you control is whether your pages are reachable, readable and correct; the technical SEO guide covers the crawl and indexing basics.
What Sitelemetry checks about llms.txt
Sitelemetry's AI visibility area is part of the six-area audit (security, technical SEO, AI visibility, accessibility, performance and integrations), which starts at $49/month on the Starter plan. Its llms.txt check is narrow:
- It requests
/llms.txtonce at the origin and reports the status and size. - The file counts as present only if it returns a 2xx status, is not empty, is not HTML and looks like a link guide: a line starting with
#or-, a URL, or words such as docs or sitemap. A missing file is reported as a partial result with a recommendation. - It does not validate the format, does not follow or test the links, and does not claim that any AI system reads the file.
The same area checks robots.txt for site-wide blocks against 16 named AI crawlers, including GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot and Google-Extended. Rules under User-agent: * and partial disallows are not attributed to individual AI crawlers (a site-wide block for all crawlers is reported separately), and the list is fixed, so check crawlers that are not on it, such as Claude-SearchBot or Claude-User, yourself. It also reports JSON-LD structured data without validating it against schema.org. The audit reads static HTML from a sample of pages without running JavaScript, and it does not query ChatGPT, Claude, Perplexity or Google or track citations in AI answers.
The Free plan covers security checks only. The free Launch Readiness Snapshot runs eight passive checks without signup; it does not look at llms.txt, robots.txt or structured data, and it is not the six-area audit. Plan details are on the pricing page.
Common questions
Is llms.txt an official standard?
No. It is an informal proposal by Jeremy Howard, first published in September 2024; version 2 is dated 10 August 2026. It is not an IETF or W3C standard, and no search engine or AI provider is required to read it.
Will llms.txt get my site cited in ChatGPT, Perplexity or Google's AI answers?
No provider has documented that it does. Google Search says it ignores llms.txt, so the file neither helps nor harms visibility there. As of 1 October 2026, OpenAI, Anthropic and Perplexity describe their crawlers in terms of robots.txt and do not say that they read a site's llms.txt.
Can llms.txt block AI crawlers or AI training?
No. It does not grant or deny access. Use robots.txt with the tokens each provider documents, such as GPTBot or ClaudeBot for training crawlers, and keep private content behind authentication, because some user-triggered fetchers may not follow robots.txt.
Is an llms.txt generator good enough?
For a first draft, yes. You still need to choose the pages that matter, write accurate notes, remove anything blocked in robots.txt or set to noindex, and regenerate the file when the docs change.
How can I check my llms.txt?
Fetch it and confirm a 200 status, Markdown rather than an HTML page, the H1 first and working links. Lighthouse's experimental audit and Sitelemetry's AI visibility area both check that the file is present; neither can tell you whether a particular assistant reads it.
Sources & further reading
- llmstxt.org: The /llms.txt file (proposal, v2)llmstxt.org
- Google Search Central: Optimizing your website for generative AI features on Google Searchdevelopers.google.com
- Chrome for Developers: Lighthouse llms.txt auditdeveloper.chrome.com
- Chrome for Developers: Lighthouse Agentic Browsing scoringdeveloper.chrome.com
- OpenAI: Overview of OpenAI crawlersdevelopers.openai.com
- Claude Help Center: Does Anthropic crawl data from the web, and how can site owners block the crawler?support.claude.com
- Perplexity: Perplexity crawlersdocs.perplexity.ai
- RFC 9309: Robots Exclusion Protocolwww.rfc-editor.org
- Google Crawling Infrastructure: Google's common crawlers (Google-Extended)developers.google.com
- Google Search Central: Learn about sitemapsdevelopers.google.com
Prepared by the Sitelemetry editorial team. Explore the linked sources for further detail.



