Key takeaways
- llms.txt is a proposal from Jeremy Howard, first published on 3 September 2024: a Markdown file at /llms.txt that gives AI agents a curated map of your site.
- Google's AI optimization guide says Google Search itself doesn't use these files, and that creating them neither harms nor helps your visibility or rankings in Google Search.
- Ahrefs checked 137,210 domains in May 2026: 97% of the llms.txt files it found received zero requests, and AI retrieval bots made up just 1.1% of the requests that did arrive.
- Nobody has published evidence that an llms.txt file improves citations in ChatGPT, Claude or Perplexity.
- It can make sense for documentation sites, where coding agents follow the links. OpenAI, Anthropic, the Gemini API and Perplexity all publish one for their own developer docs.
llms.txt is a proposed Markdown file, placed at yourdomain.com/llms.txt, that gives AI agents a curated list of the pages worth reading. Google says Google Search doesn’t use it, and we found no AI search engine that says it reads it.
The idea is tidy, and it has been widely promoted as a way to get AI engines to notice you.
The evidence doesn’t support the sales pitch. This guide covers what the file is, how to write one, who actually reads it, and what the companies that matter have said, with every figure traced to its source.
What Is llms.txt and Where Did It Come From?
llms.txt is a proposal, not an official web standard: a plain Markdown file that lists the most useful pages on your site so a language model can find them without wading through navigation, ads and scripts. It was put forward by Jeremy Howard of Answer.AI and fast.ai.
Who proposed it, and when?
According to llmstxt.org, the proposal was first published on 3 September 2024, and the site now carries a “v2” modified on 10 August 2026. Howard’s argument is that web pages are built for people, and that agents are best served by concise, expert-level information gathered in one place.
Why was it proposed?
Context windows are limited, and converting a cluttered HTML page back into clean text is imprecise. A short index of clean Markdown pages gets round both problems, at least in theory.
The proposal says it mainly expected the file to be useful for inference, meaning the moment an agent is answering a question, rather than for training. That framing matters later, because it’s the part nobody has confirmed for search.
How is it different from robots.txt and sitemaps?
They do different jobs and you can use all three. The table sets them side by side, using llmstxt.org’s own description of each.
| File | What it does | When it’s used | Controls access? |
|---|---|---|---|
| robots.txt | Tells automated tools what access is acceptable | Before a bot crawls | Asks bots to stay out |
| sitemap.xml | Lists all indexable human-readable pages | During crawling and indexing | No |
| llms.txt | Gives a curated, LLM-readable overview of a site or path | On demand, when an agent needs information | No |
If you want to control which bots can crawl you, that’s still robots.txt.
What Does an llms.txt File Look Like?
An llms.txt file is Markdown with one required element, an H1 with the site name, followed by optional sections in a fixed order. It lives at /llms.txt or at a subpath such as /docs/llms.txt, and a file covers the URLs under its path.
What are the parts, in order?
The spec lists them as an optional byte-order mark, the H1, a blockquote summary, optional detail text with no headings, then H2 sections containing lists of links. A section named “Optional” is used, by convention, for secondary links an agent can skip.
Each link in a file list is a Markdown hyperlink, optionally followed by a colon and a note. Where more than one file applies to a page, agents should use the most specific one.
A worked example
Here’s a minimal file for a made-up Cheltenham accountancy firm, written to the spec:
# Example Accountants Ltd
> Example Accountants Ltd is a Cheltenham accountancy firm serving sole traders and small limited companies with bookkeeping, tax returns and payroll.
Prices on our site are shown excluding VAT. We do not offer financial advice.
## Services
- [Bookkeeping](https://example.com/bookkeeping/): Monthly bookkeeping for sole traders and small companies
- [Self assessment](https://example.com/self-assessment/): Tax return preparation and filing
- [Payroll](https://example.com/payroll/): Payroll for businesses with up to 20 staff
## Pricing and contact
- [Fees](https://example.com/fees/): Fixed monthly fees for each service
- [Contact](https://example.com/contact/): Address, phone number and booking form
## Optional
- [News and guides](https://example.com/guides/): Articles on deadlines and tax changes
What else does v2 recommend?
Version 2 also proposes clean Markdown versions of pages at the same URL with .md added, plus link relations to help clients find them: rel="alternate" for the Markdown copy and rel="describedby" pointing at the llms.txt file that covers the page. These can sit in HTML <link> elements or an HTTP Link: header.
That’s extra work, and it’s the part Google specifically says you don’t need for Google Search. For a small business site, the single index file is the realistic scope.
Do Google and the AI Companies Use llms.txt?
Google has said plainly that Google Search doesn’t. For OpenAI, Anthropic and Perplexity, we found no statement that their assistants read the file when answering questions.
| Source | What it said | When |
|---|---|---|
| Google AI optimization guide | Google Search itself doesn’t use LLMS.txt files; making them neither harms nor helps visibility or rankings | Page last updated 10 July 2026 |
| Google documentation changelog | Added a note to the guide clarifying Google Search’s usage of llms.txt files | 15 June 2026 |
| Chrome Lighthouse documentation | Describes llms.txt as an emerging convention; the audit is marked N/A if the file is missing, as providing it is optional at the moment | Page last updated 5 May 2026 |
| John Mueller, Google | “FWIW no AI system currently uses llms.txt” | 17 June 2025 |
| Ahrefs server-log study | 97% of llms.txt files received zero requests | May 2026 data, published 15 June 2026 |
What Google says
Google’s guide to generative AI features lists LLMS.txt files under things you can ignore. It says you don’t need new machine readable files, AI text files, markup or Markdown to appear in Google Search, including its generative AI capabilities.
Google’s documentation changelog explains the 15 June 2026 note as an answer to community questions, and adds that it’s fine to keep these files for other services or systems that use them.
Earlier, on Bluesky, John Mueller wrote “FWIW no AI system currently uses llms.txt”, as Search Engine Roundtable reported. He added that server logs show consumer chatbots fetching pages for training and grounding but not the llms.txt file.
Why does Chrome check for it, then?
It looks like a contradiction, and Ahrefs noticed it too. Chrome’s Lighthouse documentation says that without the file, agents may spend more time crawling the site to understand its structure.
But the audit only flags a server error when fetching the file, and it’s marked not applicable on a 404. That’s browser agents and developer tooling, not Search ranking or AI Overviews.
Which AI companies publish their own?
We fetched these on 3 October 2026, and each returned a file: OpenAI at developers.openai.com/llms.txt, Anthropic at platform.claude.com/llms.txt, the Gemini API at ai.google.dev/gemini-api/docs/llms.txt and Perplexity at docs.perplexity.ai/llms.txt. llmstxt.org lists OpenAI, Anthropic and Gemini as publishers, and Perplexity’s file is live too.
All four are developer documentation, which is where the format makes most sense. Publishing a file to help agents read your own docs is different from saying your crawlers read everyone else’s.
Related: Should You Block AI Crawlers Like GPTBot?
How Many Sites Publish llms.txt?
Tens of thousands do, and the number is growing fast, but adoption isn’t the same as use. The counts depend on who is measuring and how.
What the trackers found
Originality.ai tracked more than 3 million websites and recorded growth from 4,088 llms.txt instances in June 2025 to 36,120 in May 2026, an 8.8-fold rise. Counting llms-full.txt and ai.txt as well, it found 38,980 sites.
Ahrefs looked at the 137,210 domains in Ahrefs Web Analytics that received traffic in May 2026 and found that 28% publish an llms.txt, about 38,000 domains. Ahrefs says its customers skew more technical and SEO-aware than the web at large, so the 28% is an upper bound.
Why adoption is outrunning use
Plenty of files aren’t a deliberate choice. llmstxt.org notes that documentation platforms generate one automatically and lists Yoast SEO and All in One SEO among WordPress plugins with a generator, and says Wix generates a file for every Wix site.
Ahrefs makes the same point: adoption has been driven by speculation that AI platforms may start consuming the file, not by any confirmation that they do.
What Does the Data Say About Whether Anyone Reads It?
Very little reads it: Ahrefs found that 97% of the llms.txt files it identified received no requests at all in May 2026. The study is the best public evidence so far, and the caveats matter.
What Ahrefs measured
Ahrefs’ study, published on 15 June 2026, checked every domain root for an llms.txt returning HTTP 200 and examined requests to the file in Ahrefs Bot Analytics. Of the roughly 38,000 domains with a valid file, 97% saw no requests in May, and the remaining 3% (about 1,100 domains and 22,000 requests) received all the traffic measured.
Of those requests, 96% came from bots. The table shows who was asking.
| Requester category | Share of requests |
|---|---|
| SEO audit tools | 21.7% |
| Other and unidentified | 14.9% |
| General web crawlers (such as Googlebot) | 13.1% |
| Tech profiling tools | 11.6% |
| AI agents and agentic infrastructure | 10.5% |
| GEO and AEO tools | 5.8% |
| AI training crawlers | 5.3% |
| llms.txt discoverability bots | 3.6% |
| Service and social bots (link previews) | 2.9% |
| Research bots | 2.7% |
| AI assistants (such as ChatGPT-User) | 2.5% |
| AI retrieval bots (such as OAI-SearchBot, PerplexityBot) | 1.1% |
The remainder, about 4%, was human.
What the numbers do and don’t prove
Ahrefs adds a caveat worth repeating: “fetched” doesn’t mean “read”, so every figure is a ceiling on actual consumption. Its 19.5% for all AI bots combined is “the most generous possible reading”.
Two more findings stand out. Chrome’s Lighthouse audit made just 22 requests, roughly 1 in 1,000, and Ahrefs saw zero requests from AI bots for llms.txt files that don’t exist.
So AI tools don’t go looking for the file. Ahrefs concludes that they fetch it when a link, an index or a user instruction tells them it exists.
What’s still unknown
Server logs show who asked for a file, not what a model knows. A system could use llms.txt without touching your server, for example by reading a cached copy, though nobody has shown that.
Equally, we found no study showing that sites with an llms.txt get cited more. That’s the claim the sales pitch relies on, and it’s missing.
Our guide to how to get your brand cited by ChatGPT covers what the evidence does support.
Does Your Site Need an llms.txt File?
Probably not for visibility: if you publish developer documentation that coding agents use it is cheap and sensible, but a marketing site hoping to be cited should spend the time elsewhere. The decision follows who your readers are.
| Type of site | Worth publishing? | Why |
|---|---|---|
| API or developer documentation | Yes | llmstxt.org says docs are where coding agents follow llms.txt most heavily |
| Product site whose customers use coding agents | Reasonable, low effort | Ahrefs says coding agents are the closest thing to an intended audience |
| Service business or content site chasing citations | Low priority | AI retrieval bots barely fetch the file, and Google ignores it |
| Site on a platform that generates one | Review it, then leave it | The cost is nearly zero, but check what it lists |
When it’s worth doing
If you run a docs site, a generated file costs almost nothing and the likely readers are coding agents. Ahrefs found Claude Code, Anthropic’s coding agent, out-fetched every AI retrieval bot, assistant and training crawler apart from GPTBot and one other agentic crawler.
When to skip it
For a service business or a content site, your effort is better spent on the things Google’s guide does recommend, such as non-commodity content and sound technical SEO. That is the day-to-day work of any AI search programme, and our piece on generative engine optimisation covers the wider picture, with proper structured data coming first.
Related: What Is Schema Markup and Which Types Matter?
If you want to see how AI engines describe your brand today, talk to us about AI search before you spend time on a file nobody has confirmed they read.
How Do You Create an llms.txt File?
Pick the pages you’d want an agent to quote, write a short summary and a linked list in Markdown, save it as llms.txt at your site root, and link to it. Keep it small, because the file itself should stay small enough to fit in context.
Step by step
- List the pages that explain what you do, what it costs and how to contact you.
- Write the H1, a one or two sentence blockquote summary, and one H2 section per group of links.
- Add a short note after each link, in plain language, with no marketing claims.
- Save it as
llms.txt, serve it as plain text or Markdown and place it at the root. - Open yourdomain.com/llms.txt in a browser and check it returns a 200 response, not a redirect or an HTML page.
How do agents find it?
Ahrefs’ point that AI tools fetch the file only when told it exists has a practical consequence: link to it. The v2 proposal recommends a rel="describedby" relation, which looks like this in the <head> of your pages:
<link rel="describedby" href="/llms.txt">
An HTTP Link: header does the same job and can be added in server or CDN configuration without editing pages.
What should stay out?
Leave out pages you wouldn’t want quoted, such as out-of-date posts, duplicate pages and thin pages. Ahrefs also recommends keeping the content to plain links and descriptions, with nothing instruction-shaped, because some researchers are already probing llms.txt as a prompt injection route.
How Do You Check Whether Anyone Reads Your llms.txt?
Search your server logs for requests to /llms.txt and group them by user agent. A request tells you the file was fetched, not that anything acted on it.
A log check
This command counts requests for the file by common bot name in a standard access log:
grep "GET /llms.txt" access.log | grep -oE "GPTBot|OAI-SearchBot|ChatGPT-User|ClaudeBot|Claude-SearchBot|Claude-User|PerplexityBot|Perplexity-User|Googlebot|bingbot|Slackbot|AhrefsBot|SemrushBot" | sort | uniq -c | sort -rn
If Ahrefs’ pattern holds for you, expect a few audit tools and link-preview bots, little else, and plenty of 404s if you haven’t published one. Ahrefs also found that traffic for a missing file came 98% from humans, presumably SEOs checking competitors.
Treat it like code
Because agents are built to trust the file, a stale or tampered one misleads every agent that reads it. Ahrefs suggests version-controlling it, restricting who can edit it, setting an alert for unauthorised changes, and reviewing anything a platform generates for you.
The simplest sanity check is whether AI engines can see and describe you at all. An AI visibility audit answers that question directly, and it doesn’t depend on any one file.
Related: How Do You Measure AI Search Visibility?
To find out what AI search engines can actually reach and say about you, see our AI search service, including an audit from £500.
- 1H1 titleThe name of the site or project. This is the only required part.
- 2Blockquote summaryA short summary with the key facts needed to understand the rest.
- 3Detail textOptional paragraphs or lists, but no headings, explaining how to read the file.
- 4H2 file listsSections of links to the pages or Markdown files that hold the detail.
- 5Optional sectionBy convention, links an agent can skip when it needs less context.
FAQs
Does Google crawl my llms.txt file?
Google's guide says it may discover, crawl and index many kinds of files, and that this doesn't mean a file is treated in a special way. Ahrefs saw Googlebot fetch llms.txt files about 900 times in May 2026 and read that as routine crawling, the same as for a sitemap or any other URL.
Can llms.txt hurt my SEO?
Google says it won't: these files neither harm nor help visibility in Google Search. The real costs are the time to build and maintain the file, and the risk that a stale or tampered file misleads any agent that reads it.
Where does the file go and what is it called?
Name it llms.txt and place it at the root, so yourdomain.com/llms.txt. The spec also allows one at a subpath such as /docs/llms.txt, which then covers the URLs under that path.
Can llms.txt block AI crawlers?
No. Ahrefs puts it plainly: despite the filename, it isn't a robots.txt-style directive and it controls nothing. Crawler access is still handled in robots.txt, a WAF or a CDN.
What is llms-full.txt?
It's a larger companion file that some documentation sites publish alongside llms.txt, holding the full text of their docs in one place. OpenAI's API docs file and Anthropic's file both link one, but we didn't find it in the v2 spec text, so treat it as a convention rather than part of the proposal.
Should I create separate Markdown pages for AI bots?
Google's guide says you don't need Markdown versions of your pages to appear in Google Search. The v2 proposal does suggest .md versions of pages at the same URL, but that is the proposal's recommendation, not something any search engine has asked for.
Will a WordPress plugin or website builder make one for me?
Often, yes. llmstxt.org lists Yoast SEO and All in One SEO as WordPress plugins with generators, and says Wix generates a file for every Wix site. Read whatever a plugin produces, because an auto-generated file can list pages you wouldn't choose.
Does ChatGPT read llms.txt?
OpenAI hasn't said so in the crawler documentation we reviewed. In Ahrefs' logs, GPTBot was the biggest AI fetcher among training crawlers, but OAI-SearchBot and PerplexityBot together made only a couple of hundred requests across thousands of sites.
Sources
- The /llms.txt file, v2, llmstxt.org (Jeremy Howard)
- Optimizing your website for generative AI features on Google Search, Google Search Central
- Google Search Central documentation updates, Google Search Central
- llms.txt (Lighthouse agentic browsing audits), Chrome for Developers
- We Analyzed 137K Sites: 97% of llms.txt Files Never Get Read, Ahrefs
- LLMs.txt Tracking Study and Live Dashboard, Originality.ai
- Google: No AI System Currently Uses LLMs.txt, Search Engine Roundtable
- OpenAI Developers llms.txt, OpenAI
- Anthropic Developer Documentation llms.txt, Anthropic
- Gemini API llms.txt, Google AI for Developers
- Perplexity llms.txt, Perplexity
Related services