What Is GEO and How to Get Cited in AI Answers
Generative engine optimization explained honestly: what works, what Google says, which tactics are myths, and a 30-day plan to get cited in AI answers.
Key takeaways
llms.txt is a plain-text file in Markdown format, placed at the root of a website (the address /llms.txt), that gives language models a summary of the site and a curated list of links to its important pages. It is a community proposal, not an official standard, and its effect on citations in ChatGPT, Perplexity or Claude has not been publicly demonstrated. Google says it does not use it.
Many people describe it as "robots.txt for AI." The comparison misleads: robots.txt controls access, while llms.txt merely suggests what to read. Below you will find the format, an example, what is known and unknown, and how to decide whether it is worth it. For the wider picture of what actually drives visibility in AI answers, see our generative engine optimization service and the overview linked later in this article.
llms.txt is a specification proposed in September 2024 by Jeremy Howard to give language models a clean "map" of a website. The motivation is practical: a model cannot read an entire site, and HTML pages are full of navigation, scripts and ads. A short Markdown file with links and descriptions would save effort and reduce errors.
The proposal also suggests something more: important pages could have a clean Markdown version at the same URL with an .md extension added or substituted. The idea fits technical documentation best, where a coding assistant can read a page directly without the clutter.
The file carries no recognized authority. It has not been standardized by the W3C or the IETF, unlike robots.txt, which is described in RFC 9309. So no platform is obliged to read it.
According to the proposal, the only required element is an H1 heading with the name of the site or project. Then, optionally:
The general structure, line by line, for a fictional business ("Example Clinic," physiotherapy in Bucharest):
| Element | What the line looks like | Role |
|---|---|---|
| Title (required) | # Example Clinic: physiotherapy in Bucharest | Name of the site or project |
| Summary | > Physiotherapy and medical rehabilitation in District 2, by appointment. | Essential context in one or two sentences |
| Details | a short paragraph of plain text | Clarifications: area served, language, what you do not offer |
| Section | ## Services | Groups links by topic |
| Link with description | a dash, then [Physiotherapy] followed immediately by the full address in parentheses, then : what we treat and how a session works | One resource and why it matters |
| Secondary section | ## Optional | Resources an agent may skip |
A real file has one section each for services, about, and contact, and every link line follows the pattern in the table. Addresses are always complete (https plus your domain), never relative.
https://your-domain.com/llms.txt as plain text in UTF-8.If your business operates in Romania, keep the descriptions in the language of the pages they point to. A Romanian page described in English confuses both people and agents.
For Google there is a public position, in its documentation on generative features: you do not need to create new machine-readable files, "AI text files," markup or Markdown to appear in Google Search, and Google Search itself does not use them. Google adds that doing so will neither help nor harm visibility or rankings in Search, and its AI features (AI Overviews, AI Mode) are part of Search.
For other engines, the situation is unclear:
So the accurate wording is: the effect on citations is unproven. Any agency promising gains "thanks to llms.txt" owes you evidence, not arguments.
| Situation | Does llms.txt help? | Note |
|---|---|---|
| Google Search, AI Overviews, AI Mode | No | Google says it does not use such files |
| Technical documentation used by coding assistants | Possibly | Clearest use case; depends on the tool |
| Online store or service site, aiming at chatbot citations | Unconfirmed | No solid public evidence |
| Your own agents (you hand them the file) | Yes | You control what your agent reads |
| Controlling bot access | No | That is what robots.txt is for |
A file is useful only if it can be read without problems. Check the following:
/llms.txt must return 200, not a redirect to the homepage and not a custom error page that looks like full HTML.noindex. A link to a blocked page contradicts the file's purpose.As of this writing there is no reliable public source showing, on a large sample and with a stated method, that adding llms.txt raises how often an AI engine cites a site. We have no such data either, and we will not invent figures. What we can say accurately:
If you test it, treat it as an experiment: record the questions, the publish date, and the cited sources before and after, and repeat the test several times. One different answer proves nothing.
You cannot find out from the file itself. You have two real sources:
/llms.txt and note the user agent and IP. Compare with the official lists of bot addresses (OpenAI, Anthropic and Perplexity each publish their IP ranges). A bot request does not mean influence over answers, only that the file was read.If you are targeting Google, put your resources elsewhere: indexing, useful content, snippet eligibility. See how to become a source in Google AI Overviews. At best, llms.txt remains an optional layer on top of these foundations.
Publish llms.txt if you have documentation or a set of resources you want to hand to agents cleanly, if you can maintain the file, and if you are not expecting miracles. Next steps: choose the essential pages, write the summary, publish the file, check your logs after a month, and record the results. If you want to decide together whether it makes sense for your site, see our GEO service page.
Frequently asked questions
The proposal was published in September 2024 by Jeremy Howard, known for his work in machine learning. It is not a standard adopted by a body such as the W3C or the IETF. It is a community specification hosted at llmstxt.org.
No. Robots.txt tells bots what they may access, a sitemap lists pages for discovery, and llms.txt is a curated list of resources with descriptions for models. The proposal itself says it complements those files rather than replacing them.
No. The point is selection: a few dozen essential pages, each with a short description. A file with thousands of links imitates a sitemap and loses its purpose. Secondary pages can go in a section labeled "Optional."
Not through a penalty: Google says it ignores such files. The risks are practical: outdated information, dead links, or content that differs from what is visible on the site, which can mislead an agent that reads it.
Check your server logs for requests to /llms.txt and look at the user agent and IP address. Compare them with the lists published by OpenAI, Anthropic or Perplexity. If you see no requests, you have no evidence of use, however good the file looks.
Related service
Keep reading
Generative engine optimization explained honestly: what works, what Google says, which tactics are myths, and a 30-day plan to get cited in AI answers.
Measuring AI search visibility: documented test questions, referral traffic in analytics, server logs, Search Console, and the limits of every method.
AI crawlers in robots.txt: what GPTBot, OAI-SearchBot, ClaudeBot and PerplexityBot do, how to separate training from search, with copy-ready rules and limits.
Send us your website address and we’ll reply with a free initial analysis and a concrete SEO strategy — no strings attached.