AI crawlers fall into three groups, and robots.txt treats them differently: training crawlers (like GPTBot and ClaudeBot) collect content for model training, search crawlers (like OAI-SearchBot, Claude-SearchBot and PerplexityBot) index pages so assistants can cite them, and user-triggered fetchers (like ChatGPT-User, Claude-User and Perplexity-User) visit a page when someone asks. Blocking GPTBot keeps your content out of OpenAI’s training. It does not remove you from ChatGPT search. Blocking OAI-SearchBot does.
Many sites block “AI bots” as one category, then disappear from AI answers. This reference lists each vendor’s documented user agents, explains what blocking each one does, and gives copy-paste robots.txt setups for the three most common policies. All names and behaviors come from the vendors’ own documentation, last checked on 9 October 2026.
Key takeaways
- Each major AI vendor separates training crawlers, search crawlers and user-triggered fetchers, and each one reads its own robots.txt rule.
- Blocking a training crawler such as GPTBot or ClaudeBot doesn’t, by itself, remove you from that vendor’s AI search results.
- Blocking a search crawler such as OAI-SearchBot or PerplexityBot can stop that assistant from showing and citing your pages.
- User-triggered fetchers may not follow robots.txt: OpenAI says rules “may not apply” to ChatGPT-User, and Perplexity says Perplexity-User “generally ignores” them.
- Google-Extended controls Gemini training and grounding, not Google Search. AI Overviews and AI Mode are governed by Googlebot rules.
Training bots vs search bots vs user-triggered fetchers
Training crawlers gather public pages that may be used to train future models. Blocking them is a training opt-out, nothing more.
Search crawlers build the index an assistant searches when it answers with web results. If one is blocked, the assistant has less or nothing of yours to retrieve, so it can’t cite your pages for prompts like “best CRM for a 20-person startup”.
User-triggered fetchers act on behalf of a person in the chat, for example when someone pastes your URL and asks for a summary. They don’t crawl the web automatically. Because a human requested the page, some vendors say robots.txt may not apply.
For the bigger picture on how retrieval shapes answers, see how AI assistants choose their sources.
The AI crawler reference table
Robots.txt tokens as documented by each vendor (last checked 9 October 2026). Version numbers in full user-agent strings change, so match on the token, not the full string.
| Vendor | Token | Type | What the vendor says it does |
|---|---|---|---|
| OpenAI | GPTBot |
Training | Crawls content that may be used to train OpenAI’s foundation models |
| OpenAI | OAI-SearchBot |
Search | Surfaces websites in ChatGPT’s search features |
| OpenAI | ChatGPT-User |
User-triggered | Certain user actions in ChatGPT and Custom GPTs; not used for automatic crawling |
| OpenAI | OAI-AdsBot |
Ads review | Visits only pages submitted as ads on ChatGPT, to check their safety |
| Anthropic | ClaudeBot |
Training | Collects web content that could contribute to model training |
| Anthropic | Claude-SearchBot |
Search | Indexes content to improve the quality of Claude’s search results |
| Anthropic | Claude-User |
User-triggered | Fetches pages when a Claude user asks a question that needs them |
| Perplexity | PerplexityBot |
Search | Surfaces and links websites in Perplexity results; not used for training foundation models |
| Perplexity | Perplexity-User |
User-triggered | Visits a page to answer a user’s question; not used for crawling or training |
Google-Extended |
Control token | Manages use of crawled content for Gemini training and grounding |
Sources: OpenAI crawlers overview, Anthropic’s crawler help article, Perplexity crawlers, and Google’s common crawlers list.
Google-Extended has no user-agent string of its own: Google crawls with its usual user agents and reads the token as a permission. It covers Gemini training and grounding in Gemini Apps and Vertex AI, so it isn’t purely a training switch.
What happens when you block each one
- GPTBot. OpenAI says disallowing it “indicates a site’s content should not be used in training generative AI foundation models”. OpenAI also states that each setting is independent, so you can stay in ChatGPT search while opting out of training.
- OAI-SearchBot. Opted-out sites aren’t shown in ChatGPT search answers, though they can still appear as navigational links. OpenAI notes that robots.txt changes can take about 24 hours to reach its search systems.
- ClaudeBot. Anthropic says blocking it signals that your future materials should be excluded from training datasets.
- Claude-SearchBot. Anthropic says blocking it prevents indexing for search, which may reduce your visibility and accuracy in Claude’s search results.
- Claude-User. Claude can no longer retrieve your pages when users ask, which may reduce visibility.
- PerplexityBot. Perplexity recommends allowing it so your site can appear in its search results.
- Google-Extended. Google says it “does not impact a site’s inclusion in Google Search nor is it used as a ranking signal”. It won’t take you out of AI Overviews or AI Mode, which use Googlebot.
If you want to limit what Google shows from your pages in Search, including AI features, Google points to nosnippet, data-nosnippet, max-snippet and noindex. Those affect classic results too, so use them with care.
Recommended robots.txt setups
Under the robots.txt standard, a crawler follows the group that names it most specifically and ignores the User-agent: * group. If you add a group for GPTBot, repeat any general Disallow rules you want it to respect inside that group.
Stay visible in AI search, opt out of training
Common for B2B sites: citable pages, no training.
User-agent: GPTBot
Disallow: /
User-agent: ClaudeBot
Disallow: /
User-agent: Google-Extended
Disallow: /
User-agent: OAI-SearchBot
Allow: /
User-agent: Claude-SearchBot
Allow: /
User-agent: PerplexityBot
Allow: /
Note the Google-Extended trade-off: it also covers grounding in Gemini Apps, so blocking it may limit how Gemini uses your content in answers.
Full opt-out
User-agent: GPTBot
Disallow: /
User-agent: OAI-SearchBot
Disallow: /
User-agent: ChatGPT-User
Disallow: /
User-agent: ClaudeBot
Disallow: /
User-agent: Claude-SearchBot
Disallow: /
User-agent: Claude-User
Disallow: /
User-agent: PerplexityBot
Disallow: /
User-agent: Perplexity-User
Disallow: /
User-agent: Google-Extended
Disallow: /
This removes you from these vendors’ AI search, not just training. User-triggered fetchers may still reach pages, per OpenAI and Perplexity, and AI Overviews stay tied to Googlebot.
Allow everything
Robots.txt allows by default. If no group names these bots and your User-agent: * group doesn’t disallow them, they can already crawl. Explicit Allow: / groups simply document your intent.
How to check your current robots.txt and server logs
- Read the live file. Open
yourdomain.com/robots.txtand search for each token in the table. Check every subdomain too, since each has its own file. - Check the
*group. A broadDisallowunderUser-agent: *also applies to any AI bot without its own group. - Search your server or CDN logs for the tokens (GPTBot, OAI-SearchBot, ClaudeBot, Claude-SearchBot, PerplexityBot and the user fetchers). Note status codes: a wall of 403s means something other than robots.txt is blocking.
- Verify the IPs. User-agent strings can be spoofed. OpenAI, Anthropic and Perplexity publish IP ranges in JSON files linked from their docs.
- Wait, then re-check. OpenAI and Perplexity both mention delays of up to about 24 hours.
Beyond robots.txt: CDNs and bot protection
A clean robots.txt doesn’t help if your firewall blocks the bot first. CDN and WAF tools can challenge automated traffic, sometimes through one “block AI bots” setting that doesn’t separate search from training.
Perplexity’s docs give Cloudflare and AWS WAF instructions for allowing its bots by user agent and published IP range. OpenAI’s documentation likewise points to its published IP lists for allowing OAI-SearchBot. Anthropic advises against blocking its bots by IP, because the bot then can’t read your robots.txt and the opt-out may not hold.
Crawlability is only a precondition. Our guides to ranking in ChatGPT and ranking in Perplexity cover what earns the citation after that.
Confirm you’re still cited with Citegram
Citegram doesn’t scan robots.txt or server logs. What it shows is the outcome: whether assistants still mention and cite you after a change.
- Pick prompts where you’re cited today, such as “best CRM for a 20-person startup”, before you edit robots.txt.
- Run them in the real apps. The Chrome extension asks each prompt in ChatGPT, Gemini, Perplexity, Claude, DeepSeek, Grok and Google AI Overviews, in background tabs, using the accounts already signed in to your browser. No API key, and questions count toward your own plan usage.
- Record the baseline. Citegram logs whether your brand is mentioned, its position, and the sources and links each answer cited.
- Re-run after the change, once the vendors’ propagation window has passed, and compare trends over repeated runs.
- Check the cited domains. If your pages drop out of citations while competitors stay, review your search crawler rules first.
Answers come from your own accounts, so personalization and location can affect them. For a fuller review, see our AI visibility audit guide.
FAQ
Should I block GPTBot?
Block GPTBot if you don’t want OpenAI to use your content for training. OpenAI documents it as independent from OAI-SearchBot, so blocking GPTBot alone doesn’t remove you from ChatGPT search. If you want to appear in ChatGPT search, keep OAI-SearchBot allowed.
Does blocking AI crawlers hurt my SEO?
Blocking AI vendors’ crawlers doesn’t affect Googlebot, and Google says Google-Extended is not a ranking signal for Google Search. What you can lose is visibility in AI assistants: blocking search crawlers like OAI-SearchBot or PerplexityBot can keep your pages out of their answers.
Does Google-Extended affect Google Search or AI Overviews?
Google says Google-Extended doesn’t affect inclusion in Google Search and isn’t a ranking signal. It manages use of your content for Gemini training and grounding. AI Overviews and AI Mode are part of Search, where Googlebot rules and snippet controls apply.
Do AI bots respect robots.txt?
Anthropic says its bots honor robots.txt directives, and OpenAI and Perplexity tell site owners to manage their search and training crawlers through robots.txt. User-triggered fetchers are the exception: OpenAI says robots.txt “may not apply” to ChatGPT-User, and Perplexity says Perplexity-User “generally ignores robots.txt rules”.
Confirm you’re still visible after changes
Changing robots.txt takes minutes. Checking that assistants still cite you is the part that matters. The free plan tracks one prompt across all major assistants, forever. Pro adds unlimited prompts for $10 a month. See pricing.


