SEO blog · GEO & AI

AI Crawlers (GPTBot, ClaudeBot, PerplexityBot): Allow or Block in robots.txt

Key takeaways

  • AI bots have different roles: training (GPTBot, ClaudeBot), search (OAI-SearchBot, Claude-SearchBot, PerplexityBot) and user-triggered requests (ChatGPT-User, Claude-User).
  • You can block training and allow search by writing separate robots.txt rules for each user agent.
  • Robots.txt is a voluntary request, not protection: some user-triggered bots (ChatGPT-User, Perplexity-User) may ignore it, and sensitive content needs authentication.
  • Google-Extended is a control token for Gemini, not a separate crawler, and it does not affect Google Search or AI Overviews.

AI crawlers are programs that access websites for products built on language models, and in robots.txt you control them by name (user agent): GPTBot and ClaudeBot for training; OAI-SearchBot, Claude-SearchBot and PerplexityBot for search; ChatGPT-User, Claude-User and Perplexity-User for requests triggered by a user. The key is not to treat all AI bots alike: you can refuse training and still stay eligible for search.

This guide explains what each bot does, how to write rules for three common scenarios, what the limits of robots.txt are, and how to verify what happens. For the strategic context, read our GEO overview first; our service is described on the generative engine optimization page.

Why are AI bots not all the same?

The three categories have different effects on your business:

  • Training bots collect content to develop models. Blocking them says, "do not use my pages to train future models." It does not remove you from a search engine.
  • Search bots build indexes for the search features of AI products. Blocking them can reduce your chances of being shown or cited in those products.
  • User-triggered bots fetch a page because a person asked for it (for example, they pasted a link into a chat). OpenAI says that because ChatGPT-User actions are initiated by a user, robots.txt rules may not apply, and Perplexity says Perplexity-User generally ignores them. Anthropic, by contrast, says Claude-User honors robots.txt.

If you mix up the categories, you make costly mistakes: you block everything "to be safe" and vanish from answers, or you allow everything without realizing you have chosen to be used for training.

Which bots exist and what do they do? A reference table

The information below comes from the providers' public documentation at the time of writing. Names and rules can change; check the official pages before you implement.

ProviderAgentRoleDoes robots.txt apply?
OpenAIGPTBotTraining generative modelsYes
OpenAIOAI-SearchBotAppearing in ChatGPT search featuresYes (takes effect in about 24 hours)
OpenAIChatGPT-UserUser-triggered visits, including custom GPTsDocumentation says rules may not apply
AnthropicClaudeBotCollecting content for modelsYes, blocked through robots.txt
AnthropicClaude-SearchBotImproving search resultsYes
AnthropicClaude-UserClaude users' requests that need web accessYes, per the support page
PerplexityPerplexityBotIndexing for Perplexity results; does not train models, per its documentationYes
PerplexityPerplexity-UserUser-triggered requestsDocumentation says it generally ignores robots.txt
GoogleGooglebotIndexing for Search, including AI OverviewsYes
GoogleGoogle-ExtendedToken for Gemini (training, grounding)A robots.txt token, not a separate crawler

OpenAI also has OAI-AdsBot, which checks pages submitted as ads and does not follow robots.txt. The documentation also notes that GPTBot and OAI-SearchBot may share crawl results if both are allowed, to avoid duplication.

A commonly misunderstood point: Google-Extended is not a crawler. It will not show up in your logs. It is a token you write in robots.txt to say whether content crawled by Google may be used to train Gemini models and for grounding in those products. According to Google's documentation, it does not affect inclusion in Search and is not a ranking signal. It also does not control AI Overviews; for those you use snippet controls, explained in our guide to Google AI Overviews.

Scenario 1: allow AI search, refuse training

This suits most businesses that want visibility without offering their content for training:

User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: Google-Extended
Disallow: /

User-agent: OAI-SearchBot
Allow: /

User-agent: Claude-SearchBot
Allow: /

User-agent: PerplexityBot
Allow: /

Each group applies only to the agent it names. If you also want to exclude user-triggered requests where they are honored (for example Claude-User), add them explicitly, knowing you will reduce the chance that a user cites your page through an assistant.

Scenario 2: block all AI bots

User-agent: GPTBot
Disallow: /

User-agent: OAI-SearchBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: Claude-SearchBot
Disallow: /

User-agent: PerplexityBot
Disallow: /

User-agent: Google-Extended
Disallow: /

The result: you leave these providers' search products, to the extent they respect the rules. You do not affect Google Search (Googlebot is not blocked). Choose this only if you have a clear business reason.

Scenario 3: block only one area of the site

User-agent: GPTBot
Disallow: /members/
Disallow: /paid-courses/

User-agent: ClaudeBot
Disallow: /members/
Disallow: /paid-courses/

Useful when you have premium content or sections that should not be used. Careful: if those pages are truly private, robots.txt is not enough; see the section on limits.

How to apply the rules, step by step

  1. Decide the policy. For each category (training, search, user), choose allowed or blocked.
  2. Write the file. You can start from our robots.txt generator and add the groups above.
  3. Keep Googlebot rules separate. Do not mix groups; each bot follows the most specific group that matches it.
  4. Publish at the root (/robots.txt), returning status 200 as plain text.
  5. Test. Open the file in a browser and check the syntax.
  6. Check your firewall and CDN. Some services block unknown bots by default; if you want to be read, you must let them in.
  7. Watch your logs for 2–4 weeks and compare them with the official IP ranges.

Verification in practice

User agents can be faked. Providers publish IP ranges for their bots (OpenAI for each agent, Anthropic through an official list, Perplexity for PerplexityBot and Perplexity-User). The minimum procedure:

  1. Extract from your logs the requests whose user agent contains "GPTBot," "OAI-SearchBot," "ClaudeBot," "PerplexityBot" and so on.
  2. Compare the IP addresses with the official lists.
  3. Flag "fake" requests (right name, wrong IP) separately; do not treat them as real bots.
  4. To test your firewall rules, make a trial request with a simulated user agent and see what response you get.

Anthropic notes a useful detail: blocking by IP address can prevent the bot from reading robots.txt, so it does not work as a "polite" refusal. If you use a firewall, do not block the official IPs of bots you want to refuse through robots.txt, so the rules can be read. Anthropic also says its bots support the non-standard Crawl-delay directive, which can help if crawling puts load on your server. Googlebot, by contrast, ignores Crawl-delay, so do not count on it for Google.

Which policy fits which type of business?

There is no universal recipe. The table below is a logical starting point, not a rule; adjust it to your goals and to legal advice where needed.

Business typeTrainingAI searchUser requestsReason
Local service businessAllow or block, your choiceAllowAllowYou want to be recommended when someone looks for a provider
Online storeYour choiceAllowAllowProduct descriptions are public and you want visibility
News or paid-content publisherBlockEvaluateEvaluateContent is the main asset; licensing matters
Course platform or members areaBlock on paid areasAllow on public pagesAllow on public pagesYou protect what you sell and expose what attracts customers
Internal or staging siteBlockBlockBlockProtect with authentication, not only robots.txt

If you are a foreign company with an audience in Romania, remember that AI engines can serve answers in Romanian based on pages in any language. Your access policy applies to the whole domain, not by language, unless you explicitly separate directories.

An example of log verification

Suppose that after publishing the Scenario 1 rules you see two kinds of requests in your server log. First: a user agent containing "OAI-SearchBot" with an IP in the range OpenAI publishes for that agent. This is a legitimate, permitted visitor; you do nothing. Second: a user agent "GPTBot" but with an address that is not in OpenAI's list. Either the lists changed since you last downloaded them, or someone is using the name falsely. In both cases, download the official list again; if the address is still missing, treat the request as unverified and, if needed, limit it at the firewall.

For each agent, also watch what it accesses: if a search bot requests huge numbers of insignificant pages (filters, parameters), the problem may lie in your site architecture, not in the bot. Our guide to pages not indexed helps you recognize such crawl traps.

The limits of robots.txt

  • It is voluntary. RFC 9309 describes the protocol as a request that well-behaved crawlers honor, not a security mechanism. Nothing stops a malicious bot.
  • User-triggered requests may ignore the rules; the OpenAI and Perplexity documentation say so for ChatGPT-User and Perplexity-User.
  • It does not erase what was already collected. A new rule applies from that moment, not retroactively.
  • It does not hide content. A blocked URL can still appear elsewhere if it is linked. For sensitive content use authentication or noindex where it applies.
  • Legal aspects of using content for training (copyright, licenses) are separate from robots.txt; talk to a lawyer, not a text file.

Common mistakes

  • One User-agent: * group for everything. General rules do not distinguish training from search.
  • Blocking Google-Extended hoping to leave AI Overviews. It does not work.
  • Blocking Googlebot by accident with a Disallow: / placed in the wrong group. Check after every change.
  • Misspelled names (for example "ChatGPT" instead of "ChatGPT-User" or "GPTBot"). Use exactly the names in the documentation.
  • Forgetting the CDN. The file is right, but the firewall rejects the bots you want to welcome.
  • No review. Providers add and rename agents; schedule a quarterly check.

When this does NOT apply or does not help

  • If your problem is plagiarism or aggressive scraping, robots.txt is not the solution; you need rate limiting, a WAF and monitoring.
  • If you want to stop a person from using an AI assistant on your page, bot rules guarantee nothing.
  • If you have no real reason to block, allowing search bots is generally the decision that keeps the visibility opportunity open.
  • If you care about Google visibility, other providers' AI bots do not affect it; classic indexing does.

How it ties into the rest of your strategy

Bot access is only the first filter. The llms.txt file does not replace access rules and has no proven effect. After you decide your policy, measure what happens: see how to measure visibility in ChatGPT and AI search. If you suspect broader accessibility problems, a technical SEO audit will surface them.

Conclusion: decide by category, verify in the logs

A sound policy for many sites is: refuse training if you want to, allow search if you want visibility, and do not rely on robots.txt for anything confidential. Next steps: pick a scenario, generate the file, check your firewall, and watch the logs for a month. If you want us to review the setup with you, see our GEO service page.

Sources and further reading

Frequently asked questions

Frequently asked questions

If I block GPTBot, do I disappear from ChatGPT?

Not necessarily. According to OpenAI, GPTBot relates to model training, while OAI-SearchBot relates to appearing in ChatGPT's search features. They are independent rules. If you block only GPTBot and leave OAI-SearchBot allowed, your content can be eligible for search but not for training.

How long does a robots.txt change take to apply?

For OAI-SearchBot, OpenAI's documentation indicates about 24 hours after a robots.txt update. The other providers we reviewed do not state a timeframe, so assume a delay. Check your server logs over the following days to see whether the bots respect the new rules.

Can I block only part of my site from AI bots?

Yes. Robots.txt supports directory rules, for example Disallow: /clients/ under a specific user agent. It works at the path level, not the paragraph level. For fragments inside a page there is no equivalent control in robots.txt; use other solutions or remove the content.

How do I know a visitor is really an OpenAI or Anthropic bot?

User agents can be spoofed. Compare the IP address with the ranges the provider publishes: OpenAI, Anthropic and Perplexity offer official lists. If the user agent says GPTBot but the IP is not on the list, you are not dealing with the real bot.

Should I block all AI bots?

It depends on your business. If you want to be recommended in answers, blocking search bots removes you from those products. If you have paid or licensed content, blocking training and adding real protection through authentication can be justified. Decide separately by bot type.

Related service

GEO — AI search optimization

See the service →

Let’s grow your site’s organic traffic

Send us your website address and we’ll reply with a free initial analysis and a concrete SEO strategy — no strings attached.