Free SEO tool · Technical · GEO
robots.txt & llms.txt generator for Google and AI crawlers
Pick your platform, decide which AI crawlers get in — search and citation only, or training too — and get a file ready to upload. Plus an llms.txt generator and a tester that applies your rules the way Google does.
A preset replaces the rules below with the usual ones for that platform. You can edit them afterward.
The summary an AI assistant reads first.
What llms.txt is: a 2024 proposal (llmstxt.org), not a standard. There is no public confirmation that the major AI engines use it systematically. It is a small investment, most useful for assistants that read a site's documentation.
Open yoursite.com/robots.txt, copy everything and paste it here. The tester doesn't download anything from the internet.
How the tester decides (RFC 9309 and Google's documentation): it picks the group with the most specific user-agent (case-insensitive) and merges groups with the same name; if none exists, it uses *. Then the rule with the longest matching path wins; on a tie, Allow wins. With no matching rule, access is allowed.
What robots.txt does, and what it doesn't
Robots.txt is a plain text file at the root of your domain that tells crawlers which parts of your site they may request. Googlebot, Bingbot and, increasingly, the bots run by AI companies read it. Used well, it saves crawl resources for the pages that matter and keeps low-value URLs out of the way: carts, customer accounts, internal search results, filter parameters.
What it doesn't do: it doesn't remove pages from Google (a blocked page can still be indexed without its content if other pages link to it), and it doesn't protect private data, because the file is public. To keep a page out of the index, use noindex; for confidentiality, use authentication. We cover the difference in our technical SEO service.
Training or search: which AI bots to allow
AI companies run different bots for different jobs. Training bots (GPTBot, ClaudeBot, CCBot) gather pages for future models: they send no traffic, though your content may shape what a model “knows”. Search bots (OAI-SearchBot, Claude-SearchBot, PerplexityBot) look things up in real time when someone asks a question and can cite your site as a linked source. User-triggered agents (ChatGPT-User, Claude-User, Perplexity-User) open a specific page on request.
For most businesses that want customers from AI search, the safe choice is to allow at least the search bots. Blocking training is a legitimate business decision, for publishers or rights-holders for instance, but make it deliberately. Our guide to AI crawlers and robots.txt explains the trade-offs, and the GEO optimization service covers how to get cited.
How Google decides which rule applies
- Only one group applies: the one with the most specific user-agent. If a Googlebot group exists, the
*rules are ignored entirely for Googlebot. - The longest match wins:
Allow: /wp-admin/admin-ajax.phpbeatsDisallow: /wp-admin/because it is more specific. On equal length, Allow wins. - Special characters:
*means anything,$anchors the end of the URL.Disallow: /*.pdf$blocks PDFs but not/guide.pdf?v=2. - Paths are case-sensitive:
/Admin/and/admin/are different rules.
Common mistakes
- A forgotten
Disallow: /after launch: the staging site was blocked, and the file went live that way. - Blocked CSS and JavaScript (for example
/wp-content/), which stops Google from rendering pages properly. - A
Noindexdirective in robots.txt: Google stopped honoring it in 2019. - Store filters left open: thousands of parameter combinations eat your crawl budget. See ecommerce SEO.
Where llms.txt fits
The llms.txt file doesn't replace robots.txt. Robots.txt says what crawlers may access; llms.txt offers, in Markdown, a summary of the site and a list of its essential pages. It is a proposal (llmstxt.org), not a standard, so treat it as a cheap bonus. Write the description like a direct answer (who you are, what you offer, where) and link only to pages worth reading. Our llms.txt guide goes deeper.
Once the files are live, test your key URLs with the tester above, make sure your XML sitemap is in order, and round out the picture with structured data and a GEO readiness score for your site.
Frequently asked questions
FAQ: robots.txt, AI crawlers and llms.txt
If I block GPTBot, can I still show up in ChatGPT?
In principle, yes. According to OpenAI’s documentation, GPTBot collects pages for model training, while ChatGPT search uses OAI-SearchBot (check each company’s current bot documentation, as these details change). If you want to be cited as a source, keep OAI-SearchBot (and ChatGPT-User) allowed and block GPTBot separately if you wish. Anthropic follows the same logic: ClaudeBot for training, Claude-SearchBot for search.
Does blocking Google-Extended hurt my Google rankings?
According to Google, no. Google-Extended isn't a separate crawler but a token that controls whether your pages can be used for Gemini model training and grounding. Google states that it doesn't affect inclusion or ranking in Google Search; Googlebot still crawls your pages.
Does robots.txt hide a page from Google?
Not reliably. Robots.txt controls crawling, not indexing: a blocked page can still show up in results, without a description, if other pages link to it. To keep a page out of the index, use a noindex meta robots tag and leave the page crawlable so the tag can be read. More in our technical SEO service.
What is llms.txt, and does it matter?
llms.txt is a 2024 proposal (llmstxt.org): a Markdown file at the root of your site that summarizes what the site does and lists the key pages. It isn't a standard, and there is no public confirmation that the major AI engines use it systematically. It takes minutes to make, so it is a small bet, but it doesn't replace good content or crawler access. See our llms.txt guide and GEO optimization.
Where do I upload the files, and how do I check they work?
Both live at the root of your domain: example.com/robots.txt and example.com/llms.txt; each subdomain needs its own robots.txt. After uploading, check the robots.txt report in Google Search Console and test your key URLs with the tester above. An SEO audit also reveals which important pages are blocked by mistake.
SEO Lab
More free tools
Core Web Vitals speed test
Measure LCP, INP, CLS and Lighthouse scores for any site, with real-user data from Google.
Open the tool →Schema markup generator
Ready-to-paste structured data: organization, local business, FAQ, article, product, breadcrumb.
Open the tool →GEO readiness score
A 2-minute diagnostic: how ready is your site to be cited by ChatGPT, Perplexity and Google AI?
Open the tool →Not sure which bots to let in?
We analyze your site and set an access strategy for Google and AI search engines together: what gets indexed, what gets cited and what stays private.
[email protected] · +40 771 430 955