Free SEO tool · Technical · GEO

robots.txt & llms.txt generator for Google and AI crawlers

Pick your platform, decide which AI crawlers get in — search and citation only, or training too — and get a file ready to upload. Plus an llms.txt generator and a tester that applies your rules the way Google does.

16 AI bots explainedRFC 9309 tester follows Google's rules3 in 1 robots · llms · tester
Crawler control · robots.txt & llms.txt Runs locally, in your browser
Site platform

A preset replaces the rules below with the usual ones for that platform. You can edit them afterward.

Rules for all robots User-agent: *

* matches any string of characters and $ marks the end of the URL. Example: /*.pdf$ blocks every PDF file. Paths are case-sensitive.

seconds · ignored by Google, honored by Bing
Crawlers and AI bots

Training means your pages can end up in the data used to build models; it doesn't bring visits. Search & citation bots look things up in real time and can show your site as a linked source. For visibility in ChatGPT, Claude or Perplexity, keep at least the search bots allowed.

Honest note: bot names are the ones the companies publish. Robots.txt is a convention honored by well-behaved crawlers, not a technical barrier, so use authentication for private content. Some companies state that agents triggered by a user request may not apply robots.txt.

Generated robots.txt

            
Checks

    What robots.txt does, and what it doesn't

    Robots.txt is a plain text file at the root of your domain that tells crawlers which parts of your site they may request. Googlebot, Bingbot and, increasingly, the bots run by AI companies read it. Used well, it saves crawl resources for the pages that matter and keeps low-value URLs out of the way: carts, customer accounts, internal search results, filter parameters.

    What it doesn't do: it doesn't remove pages from Google (a blocked page can still be indexed without its content if other pages link to it), and it doesn't protect private data, because the file is public. To keep a page out of the index, use noindex; for confidentiality, use authentication. We cover the difference in our technical SEO service.

    Training or search: which AI bots to allow

    AI companies run different bots for different jobs. Training bots (GPTBot, ClaudeBot, CCBot) gather pages for future models: they send no traffic, though your content may shape what a model “knows”. Search bots (OAI-SearchBot, Claude-SearchBot, PerplexityBot) look things up in real time when someone asks a question and can cite your site as a linked source. User-triggered agents (ChatGPT-User, Claude-User, Perplexity-User) open a specific page on request.

    For most businesses that want customers from AI search, the safe choice is to allow at least the search bots. Blocking training is a legitimate business decision, for publishers or rights-holders for instance, but make it deliberately. Our guide to AI crawlers and robots.txt explains the trade-offs, and the GEO optimization service covers how to get cited.

    How Google decides which rule applies

    • Only one group applies: the one with the most specific user-agent. If a Googlebot group exists, the * rules are ignored entirely for Googlebot.
    • The longest match wins: Allow: /wp-admin/admin-ajax.php beats Disallow: /wp-admin/ because it is more specific. On equal length, Allow wins.
    • Special characters: * means anything, $ anchors the end of the URL. Disallow: /*.pdf$ blocks PDFs but not /guide.pdf?v=2.
    • Paths are case-sensitive: /Admin/ and /admin/ are different rules.

    Common mistakes

    • A forgotten Disallow: / after launch: the staging site was blocked, and the file went live that way.
    • Blocked CSS and JavaScript (for example /wp-content/), which stops Google from rendering pages properly.
    • A Noindex directive in robots.txt: Google stopped honoring it in 2019.
    • Store filters left open: thousands of parameter combinations eat your crawl budget. See ecommerce SEO.

    Where llms.txt fits

    The llms.txt file doesn't replace robots.txt. Robots.txt says what crawlers may access; llms.txt offers, in Markdown, a summary of the site and a list of its essential pages. It is a proposal (llmstxt.org), not a standard, so treat it as a cheap bonus. Write the description like a direct answer (who you are, what you offer, where) and link only to pages worth reading. Our llms.txt guide goes deeper.

    Once the files are live, test your key URLs with the tester above, make sure your XML sitemap is in order, and round out the picture with structured data and a GEO readiness score for your site.

    Frequently asked questions

    FAQ: robots.txt, AI crawlers and llms.txt

    If I block GPTBot, can I still show up in ChatGPT?

    In principle, yes. According to OpenAI’s documentation, GPTBot collects pages for model training, while ChatGPT search uses OAI-SearchBot (check each company’s current bot documentation, as these details change). If you want to be cited as a source, keep OAI-SearchBot (and ChatGPT-User) allowed and block GPTBot separately if you wish. Anthropic follows the same logic: ClaudeBot for training, Claude-SearchBot for search.

    Does blocking Google-Extended hurt my Google rankings?

    According to Google, no. Google-Extended isn't a separate crawler but a token that controls whether your pages can be used for Gemini model training and grounding. Google states that it doesn't affect inclusion or ranking in Google Search; Googlebot still crawls your pages.

    Does robots.txt hide a page from Google?

    Not reliably. Robots.txt controls crawling, not indexing: a blocked page can still show up in results, without a description, if other pages link to it. To keep a page out of the index, use a noindex meta robots tag and leave the page crawlable so the tag can be read. More in our technical SEO service.

    What is llms.txt, and does it matter?

    llms.txt is a 2024 proposal (llmstxt.org): a Markdown file at the root of your site that summarizes what the site does and lists the key pages. It isn't a standard, and there is no public confirmation that the major AI engines use it systematically. It takes minutes to make, so it is a small bet, but it doesn't replace good content or crawler access. See our llms.txt guide and GEO optimization.

    Where do I upload the files, and how do I check they work?

    Both live at the root of your domain: example.com/robots.txt and example.com/llms.txt; each subdomain needs its own robots.txt. After uploading, check the robots.txt report in Google Search Console and test your key URLs with the tester above. An SEO audit also reveals which important pages are blocked by mistake.

    Not sure which bots to let in?

    We analyze your site and set an access strategy for Google and AI search engines together: what gets indexed, what gets cited and what stays private.