SEO blog · GEO & AI

llms.txt: What It Is, What It Does and How to Create One

Key takeaways

  • llms.txt is a community proposal: a Markdown file at your site root with a summary and links to key pages, written for language models.
  • Google says its Search ignores such files, so by Google's own account they neither help nor hurt there, including in AI Overviews.
  • There is no solid public evidence that llms.txt increases citations in ChatGPT, Perplexity or Claude; its clearest real use is technical documentation.
  • The cost is low, so expectations should be low too: publish it only if you can maintain it, and do not credit it with effects others promise.

llms.txt is a plain-text file in Markdown format, placed at the root of a website (the address /llms.txt), that gives language models a summary of the site and a curated list of links to its important pages. It is a community proposal, not an official standard, and its effect on citations in ChatGPT, Perplexity or Claude has not been publicly demonstrated. Google says it does not use it.

Many people describe it as "robots.txt for AI." The comparison misleads: robots.txt controls access, while llms.txt merely suggests what to read. Below you will find the format, an example, what is known and unknown, and how to decide whether it is worth it. For the wider picture of what actually drives visibility in AI answers, see our generative engine optimization service and the overview linked later in this article.

What is llms.txt and where does it come from?

llms.txt is a specification proposed in September 2024 by Jeremy Howard to give language models a clean "map" of a website. The motivation is practical: a model cannot read an entire site, and HTML pages are full of navigation, scripts and ads. A short Markdown file with links and descriptions would save effort and reduce errors.

The proposal also suggests something more: important pages could have a clean Markdown version at the same URL with an .md extension added or substituted. The idea fits technical documentation best, where a coding assistant can read a page directly without the clutter.

The file carries no recognized authority. It has not been standardized by the W3C or the IETF, unlike robots.txt, which is described in RFC 9309. So no platform is obliged to read it.

What does the official format look like?

According to the proposal, the only required element is an H1 heading with the name of the site or project. Then, optionally:

  • a blockquote with a short summary;
  • paragraphs with details on how to interpret the rest;
  • H2 sections, each with a list of Markdown links, optionally followed by a description after a colon;
  • a section named "Optional," which by convention holds secondary resources an agent can skip when it needs a shorter context.

The general structure, line by line, for a fictional business ("Example Clinic," physiotherapy in Bucharest):

ElementWhat the line looks likeRole
Title (required)# Example Clinic: physiotherapy in BucharestName of the site or project
Summary> Physiotherapy and medical rehabilitation in District 2, by appointment.Essential context in one or two sentences
Detailsa short paragraph of plain textClarifications: area served, language, what you do not offer
Section## ServicesGroups links by topic
Link with descriptiona dash, then [Physiotherapy] followed immediately by the full address in parentheses, then : what we treat and how a session worksOne resource and why it matters
Secondary section## OptionalResources an agent may skip

A real file has one section each for services, about, and contact, and every link line follows the pattern in the table. Addresses are always complete (https plus your domain), never relative.

How to create an llms.txt file, step by step

  1. Choose the pages. Pick 15–40 pages that explain what you do, for whom, on what terms or prices, and how to reach you. Exclude login pages, carts, tags and archives.
  2. Write the summary. Two factual sentences, no superlatives: what the company is, what it offers, where it operates.
  3. Group into sections. Services, guides, about, contact. Each link gets a description of 8–20 words that says what is on the page.
  4. Publish at the root. The file must respond at https://your-domain.com/llms.txt as plain text in UTF-8.
  5. Verify. Open it in a browser and test every link (no redirect chains, no 404s).
  6. Set a review date. Any change to services or prices means updating the file.

If your business operates in Romania, keep the descriptions in the language of the pages they point to. A Romanian page described in English confuses both people and agents.

What Google says, and what is known about other engines

For Google there is a public position, in its documentation on generative features: you do not need to create new machine-readable files, "AI text files," markup or Markdown to appear in Google Search, and Google Search itself does not use them. Google adds that doing so will neither help nor harm visibility or rankings in Search, and its AI features (AI Overviews, AI Mode) are part of Search.

For other engines, the situation is unclear:

  • the proposal itself notes that many sites publish such files, that some documentation platforms generate them automatically, and that major AI labs have files of their own for their developer documentation;
  • that shows the file exists and that some find it useful for technical use, not that ChatGPT or Perplexity use it to decide which source to cite;
  • OpenAI, Anthropic and Perplexity do not say in their bot documentation that llms.txt influences source selection.

So the accurate wording is: the effect on citations is unproven. Any agency promising gains "thanks to llms.txt" owes you evidence, not arguments.

Table: where it helps and where it does not

SituationDoes llms.txt help?Note
Google Search, AI Overviews, AI ModeNoGoogle says it does not use such files
Technical documentation used by coding assistantsPossiblyClearest use case; depends on the tool
Online store or service site, aiming at chatbot citationsUnconfirmedNo solid public evidence
Your own agents (you hand them the file)YesYou control what your agent reads
Controlling bot accessNoThat is what robots.txt is for

Common mistakes

  • Listing every URL. The file becomes a poor sitemap. For discovery, use an XML sitemap.
  • Promotional descriptions. "The best agency" informs nobody. Write what the page does.
  • Content that differs from what is visible. If the file says something different from the page, you risk misleading agents and people. Keep them aligned.
  • Dead or redirected links. An unmaintained file does more harm than none.
  • Using it as a blocking tool. Directives like "do not use my content" are not part of the proposal. For access, use robots.txt and read our guide to AI crawlers.
  • Inflated expectations. If you expect "visibility gains," you will be disappointed. If you treat it as clean documentation, it can be useful.

Technical details that decide whether the file is useful

A file is useful only if it can be read without problems. Check the following:

  • Status code. /llms.txt must return 200, not a redirect to the homepage and not a custom error page that looks like full HTML.
  • Bot access. If your firewall or CDN blocks unknown agents, the file will not be read either. Check those rules, not just robots.txt.
  • Encoding. Use UTF-8 so accented characters in descriptions display correctly.
  • Alignment with the rest. The linked pages should be in your sitemap, indexable and free of noindex. A link to a blocked page contradicts the file's purpose.
  • Markdown versions (optional). If you want to follow the proposal all the way, offer a clean Markdown version of key pages. Do it only if you can generate it automatically from the same content; two hand-maintained texts drift apart.

What we do not know, and why it matters

As of this writing there is no reliable public source showing, on a large sample and with a stated method, that adding llms.txt raises how often an AI engine cites a site. We have no such data either, and we will not invent figures. What we can say accurately:

  • the file is simple to produce and cheap to maintain;
  • it is used and appreciated in technical documentation;
  • for Google, it is irrelevant, according to official documentation;
  • for the rest, it remains a hypothesis to test, not a fact.

If you test it, treat it as an experiment: record the questions, the publish date, and the cited sources before and after, and repeat the test several times. One different answer proves nothing.

When it is NOT worth the effort

  • If the site lacks clear, indexed pages, fix those first. A summary on top of a neglected site solves nothing.
  • If nobody can maintain it, do not publish it. A two-year-old file is a liability.
  • If your only goal is Google, it is wasted effort, as Google's own documentation says.
  • If you want evidence-based decisions, move priority to verifiable work: indexing, original content, consistent brand data. The big picture is in our GEO overview.

How to check whether anything happens

You cannot find out from the file itself. You have two real sources:

  1. Server logs. Look for requests to /llms.txt and note the user agent and IP. Compare with the official lists of bot addresses (OpenAI, Anthropic and Perplexity each publish their IP ranges). A bot request does not mean influence over answers, only that the file was read.
  2. Repeated manual tests. The same questions before and after publishing, with dates. Because answers vary, a small change proves nothing. The method is in our guide to measuring visibility in ChatGPT and AI search.

How it relates to AI Overviews and the rest of your strategy

If you are targeting Google, put your resources elsewhere: indexing, useful content, snippet eligibility. See how to become a source in Google AI Overviews. At best, llms.txt remains an optional layer on top of these foundations.

Conclusion: a small file, no big promises

Publish llms.txt if you have documentation or a set of resources you want to hand to agents cleanly, if you can maintain the file, and if you are not expecting miracles. Next steps: choose the essential pages, write the summary, publish the file, check your logs after a month, and record the results. If you want to decide together whether it makes sense for your site, see our GEO service page.

Sources and further reading

Frequently asked questions

Frequently asked questions

Who created llms.txt?

The proposal was published in September 2024 by Jeremy Howard, known for his work in machine learning. It is not a standard adopted by a body such as the W3C or the IETF. It is a community specification hosted at llmstxt.org.

Does llms.txt replace robots.txt or my sitemap?

No. Robots.txt tells bots what they may access, a sitemap lists pages for discovery, and llms.txt is a curated list of resources with descriptions for models. The proposal itself says it complements those files rather than replacing them.

Should llms.txt list every page on my site?

No. The point is selection: a few dozen essential pages, each with a short description. A file with thousands of links imitates a sitemap and loses its purpose. Secondary pages can go in a section labeled "Optional."

Can llms.txt hurt my site?

Not through a penalty: Google says it ignores such files. The risks are practical: outdated information, dead links, or content that differs from what is visible on the site, which can mislead an agent that reads it.

How do I know whether anyone reads my llms.txt?

Check your server logs for requests to /llms.txt and look at the user agent and IP address. Compare them with the lists published by OpenAI, Anthropic or Perplexity. If you see no requests, you have no evidence of use, however good the file looks.

Related service

GEO — AI search optimization

See the service →

Let’s grow your site’s organic traffic

Send us your website address and we’ll reply with a free initial analysis and a concrete SEO strategy — no strings attached.