Google AI Overviews: How to Become a Cited Source
AI Overviews and AI Mode: what Google requires for a page to be a supporting link, which controls you have, how to track it in Search Console, and the limits.
Key takeaways
AI crawlers are programs that access websites for products built on language models, and in robots.txt you control them by name (user agent): GPTBot and ClaudeBot for training; OAI-SearchBot, Claude-SearchBot and PerplexityBot for search; ChatGPT-User, Claude-User and Perplexity-User for requests triggered by a user. The key is not to treat all AI bots alike: you can refuse training and still stay eligible for search.
This guide explains what each bot does, how to write rules for three common scenarios, what the limits of robots.txt are, and how to verify what happens. For the strategic context, read our GEO overview first; our service is described on the generative engine optimization page.
The three categories have different effects on your business:
If you mix up the categories, you make costly mistakes: you block everything "to be safe" and vanish from answers, or you allow everything without realizing you have chosen to be used for training.
The information below comes from the providers' public documentation at the time of writing. Names and rules can change; check the official pages before you implement.
| Provider | Agent | Role | Does robots.txt apply? |
|---|---|---|---|
| OpenAI | GPTBot | Training generative models | Yes |
| OpenAI | OAI-SearchBot | Appearing in ChatGPT search features | Yes (takes effect in about 24 hours) |
| OpenAI | ChatGPT-User | User-triggered visits, including custom GPTs | Documentation says rules may not apply |
| Anthropic | ClaudeBot | Collecting content for models | Yes, blocked through robots.txt |
| Anthropic | Claude-SearchBot | Improving search results | Yes |
| Anthropic | Claude-User | Claude users' requests that need web access | Yes, per the support page |
| Perplexity | PerplexityBot | Indexing for Perplexity results; does not train models, per its documentation | Yes |
| Perplexity | Perplexity-User | User-triggered requests | Documentation says it generally ignores robots.txt |
| Googlebot | Indexing for Search, including AI Overviews | Yes | |
| Google-Extended | Token for Gemini (training, grounding) | A robots.txt token, not a separate crawler |
OpenAI also has OAI-AdsBot, which checks pages submitted as ads and does not follow robots.txt. The documentation also notes that GPTBot and OAI-SearchBot may share crawl results if both are allowed, to avoid duplication.
A commonly misunderstood point: Google-Extended is not a crawler. It will not show up in your logs. It is a token you write in robots.txt to say whether content crawled by Google may be used to train Gemini models and for grounding in those products. According to Google's documentation, it does not affect inclusion in Search and is not a ranking signal. It also does not control AI Overviews; for those you use snippet controls, explained in our guide to Google AI Overviews.
This suits most businesses that want visibility without offering their content for training:
User-agent: GPTBot
Disallow: /
User-agent: ClaudeBot
Disallow: /
User-agent: Google-Extended
Disallow: /
User-agent: OAI-SearchBot
Allow: /
User-agent: Claude-SearchBot
Allow: /
User-agent: PerplexityBot
Allow: /Each group applies only to the agent it names. If you also want to exclude user-triggered requests where they are honored (for example Claude-User), add them explicitly, knowing you will reduce the chance that a user cites your page through an assistant.
User-agent: GPTBot
Disallow: /
User-agent: OAI-SearchBot
Disallow: /
User-agent: ClaudeBot
Disallow: /
User-agent: Claude-SearchBot
Disallow: /
User-agent: PerplexityBot
Disallow: /
User-agent: Google-Extended
Disallow: /The result: you leave these providers' search products, to the extent they respect the rules. You do not affect Google Search (Googlebot is not blocked). Choose this only if you have a clear business reason.
User-agent: GPTBot
Disallow: /members/
Disallow: /paid-courses/
User-agent: ClaudeBot
Disallow: /members/
Disallow: /paid-courses/Useful when you have premium content or sections that should not be used. Careful: if those pages are truly private, robots.txt is not enough; see the section on limits.
/robots.txt), returning status 200 as plain text.User agents can be faked. Providers publish IP ranges for their bots (OpenAI for each agent, Anthropic through an official list, Perplexity for PerplexityBot and Perplexity-User). The minimum procedure:
Anthropic notes a useful detail: blocking by IP address can prevent the bot from reading robots.txt, so it does not work as a "polite" refusal. If you use a firewall, do not block the official IPs of bots you want to refuse through robots.txt, so the rules can be read. Anthropic also says its bots support the non-standard Crawl-delay directive, which can help if crawling puts load on your server. Googlebot, by contrast, ignores Crawl-delay, so do not count on it for Google.
There is no universal recipe. The table below is a logical starting point, not a rule; adjust it to your goals and to legal advice where needed.
| Business type | Training | AI search | User requests | Reason |
|---|---|---|---|---|
| Local service business | Allow or block, your choice | Allow | Allow | You want to be recommended when someone looks for a provider |
| Online store | Your choice | Allow | Allow | Product descriptions are public and you want visibility |
| News or paid-content publisher | Block | Evaluate | Evaluate | Content is the main asset; licensing matters |
| Course platform or members area | Block on paid areas | Allow on public pages | Allow on public pages | You protect what you sell and expose what attracts customers |
| Internal or staging site | Block | Block | Block | Protect with authentication, not only robots.txt |
If you are a foreign company with an audience in Romania, remember that AI engines can serve answers in Romanian based on pages in any language. Your access policy applies to the whole domain, not by language, unless you explicitly separate directories.
Suppose that after publishing the Scenario 1 rules you see two kinds of requests in your server log. First: a user agent containing "OAI-SearchBot" with an IP in the range OpenAI publishes for that agent. This is a legitimate, permitted visitor; you do nothing. Second: a user agent "GPTBot" but with an address that is not in OpenAI's list. Either the lists changed since you last downloaded them, or someone is using the name falsely. In both cases, download the official list again; if the address is still missing, treat the request as unverified and, if needed, limit it at the firewall.
For each agent, also watch what it accesses: if a search bot requests huge numbers of insignificant pages (filters, parameters), the problem may lie in your site architecture, not in the bot. Our guide to pages not indexed helps you recognize such crawl traps.
noindex where it applies.User-agent: * group for everything. General rules do not distinguish training from search.Disallow: / placed in the wrong group. Check after every change.Bot access is only the first filter. The llms.txt file does not replace access rules and has no proven effect. After you decide your policy, measure what happens: see how to measure visibility in ChatGPT and AI search. If you suspect broader accessibility problems, a technical SEO audit will surface them.
A sound policy for many sites is: refuse training if you want to, allow search if you want visibility, and do not rely on robots.txt for anything confidential. Next steps: pick a scenario, generate the file, check your firewall, and watch the logs for a month. If you want us to review the setup with you, see our GEO service page.
Frequently asked questions
Not necessarily. According to OpenAI, GPTBot relates to model training, while OAI-SearchBot relates to appearing in ChatGPT's search features. They are independent rules. If you block only GPTBot and leave OAI-SearchBot allowed, your content can be eligible for search but not for training.
For OAI-SearchBot, OpenAI's documentation indicates about 24 hours after a robots.txt update. The other providers we reviewed do not state a timeframe, so assume a delay. Check your server logs over the following days to see whether the bots respect the new rules.
Yes. Robots.txt supports directory rules, for example Disallow: /clients/ under a specific user agent. It works at the path level, not the paragraph level. For fragments inside a page there is no equivalent control in robots.txt; use other solutions or remove the content.
User agents can be spoofed. Compare the IP address with the ranges the provider publishes: OpenAI, Anthropic and Perplexity offer official lists. If the user agent says GPTBot but the IP is not on the list, you are not dealing with the real bot.
It depends on your business. If you want to be recommended in answers, blocking search bots removes you from those products. If you have paid or licensed content, blocking training and adding real protection through authentication can be justified. Decide separately by bot type.
Related service
Keep reading
AI Overviews and AI Mode: what Google requires for a page to be a supporting link, which controls you have, how to track it in Search Console, and the limits.
llms.txt explained without hype: format, a full example, what Google says, what is known about its effect, and when it is or is not worth publishing.
Generative engine optimization explained honestly: what works, what Google says, which tactics are myths, and a 30-day plan to get cited in AI answers.
Send us your website address and we’ll reply with a free initial analysis and a concrete SEO strategy — no strings attached.