Free · No sign-up · Runs in your browser
AI robots.txt generator and checker
Choose which AI crawlers may read your site, from GPTBot to PerplexityBot, and get a robots.txt block you can paste. Or check an existing file to see which AI bots it allows or blocks, and which rule decides.
- GPTBotTraining
Content that may be used to train OpenAI's foundation models
- OAI-SearchBotAI search
Surfaces websites in ChatGPT's search features
- ChatGPT-UserUser-triggered
Certain user actions in ChatGPT and Custom GPTs, not automatic crawling
- OAI-AdsBotAds review
Only visits pages submitted as ads on ChatGPT, to check their safety
Anthropic
Vendor docs(opens in a new tab)- ClaudeBotTraining
Web content that could contribute to model training
- Claude-SearchBotAI search
Indexes content to improve Claude's search results
- Claude-UserUser-triggered
Fetches pages when a Claude user asks a question that needs them
Perplexity
Vendor docs(opens in a new tab)- PerplexityBotAI search
Surfaces and links websites in Perplexity results, not used for training
- Perplexity-UserUser-triggered
Visits a page to answer a user's question, not used for crawling or training
- Google-ExtendedGemini training & grounding
Manages use of your content for Gemini training and grounding, not Google Search
Paste it to keep your other rules. The AI section is added at the end, or replaced if it was generated here.
robots.txt
# AI crawlers (generated by citegram.com/tools/ai-robots-txt-generator) # OpenAI User-agent: GPTBot Disallow: / User-agent: OAI-SearchBot Allow: / User-agent: ChatGPT-User Allow: / User-agent: OAI-AdsBot Allow: / # Anthropic User-agent: ClaudeBot Disallow: / User-agent: Claude-SearchBot Allow: / User-agent: Claude-User Allow: / # Perplexity User-agent: PerplexityBot Allow: / User-agent: Perplexity-User Allow: / # Google User-agent: Google-Extended Disallow: / # End of AI crawlers
- Blocking Google-Extended opts you out of Gemini training and grounding. Google says it doesn't affect Google Search, and AI Overviews follow Googlebot rules.
Open yourdomain.com/robots.txt and paste its content. An example is loaded.
Summary for your home page (/)
- ChatGPT search (OAI-SearchBot), Claude search (Claude-SearchBot), and Perplexity (PerplexityBot) can read your site.
- GPTBot (OpenAI training) and ClaudeBot (Anthropic training) are blocked.
- Google-Extended is blocked: Gemini training and grounding opt-out. Google Search and AI Overviews are not affected.
- User-triggered fetchers (ChatGPT-User, Claude-User, and Perplexity-User) can open your pages when someone asks.
Result per AI crawler
OpenAI
-
GPTBotTrainingBlockedLine 9:
Disallow: /in its own group. -
OAI-SearchBotAI searchAllowed, 2 paths restrictedNo group names OAI-SearchBot, so
User-agent: *(line 2) applies, and no rule there matches /. -
ChatGPT-UserUser-triggeredAllowed, 2 paths restrictedNo group names ChatGPT-User, so
User-agent: *(line 2) applies, and no rule there matches /. -
OAI-AdsBotAds reviewAllowed, 2 paths restrictedNo group names OAI-AdsBot, so
User-agent: *(line 2) applies, and no rule there matches /.
Anthropic
-
ClaudeBotTrainingBlockedLine 13:
Disallow: /in its own group. -
Claude-SearchBotAI searchAllowed, 2 paths restrictedNo group names Claude-SearchBot, so
User-agent: *(line 2) applies, and no rule there matches /. -
Claude-UserUser-triggeredAllowed, 2 paths restrictedNo group names Claude-User, so
User-agent: *(line 2) applies, and no rule there matches /.
Perplexity
-
PerplexityBotAI searchAllowedIts group
User-agent: PerplexityBot(line 15) has no rule that matches /, so it is allowed by default. -
Perplexity-UserUser-triggeredAllowed, 2 paths restrictedNo group names Perplexity-User, so
User-agent: *(line 2) applies, and no rule there matches /.
-
Google-ExtendedGemini training & groundingBlockedLine 13:
Disallow: /in its own group.
robots.txt only decides who may read your site. Citegram shows whether assistants actually cite you.
Track this in ChatGPT, Gemini, Perplexity… with Citegram (free)How to use the AI robots.txt generator
To generate rules: start from a preset. Allow AI search, block training suits most B2B sites that want to be cited by assistants without feeding model training. Then switch any crawler between Allow and Block. If you paste your current robots.txt in the optional field, your other rules are kept and the AI section is added at the end (or replaced, if you generated it here before). Add your sitemap URL if you want it listed. Copy the result or download it as robots.txt, then upload it to the root of your domain.
To check a file: open the Check tab and paste the content of yourdomain.com/robots.txt. For each documented AI crawler, the tool shows whether it may read your home page (/), the exact line that decided it, and a short summary in plain words. A realistic example is loaded so you can see the output before pasting your own.
Everything runs in your browser. Nothing you paste is sent anywhere.
Training bots, search bots and user-triggered fetchers
Each vendor separates its crawlers by purpose, and each one reads its own robots.txt group. Training crawlers (GPTBot, ClaudeBot) collect pages that may be used to train models. Search crawlers (OAI-SearchBot, Claude-SearchBot, PerplexityBot) build the index assistants search before they cite a source. User-triggered fetchers (ChatGPT-User, Claude-User, Perplexity-User) open a page when someone asks about it.
That's why blocking GPTBot keeps your content out of OpenAI's training but doesn't remove you from ChatGPT search, while blocking OAI-SearchBot can. Google-Extended is a control token, not a crawler: Google says it manages use of your content for Gemini training and grounding and doesn't affect Google Search. AI Overviews follow Googlebot rules. The full reference, with each vendor's documentation, is in our guide to AI crawlers and robots.txt.
How the checker reads your robots.txt
The checker follows the robots.txt standard (RFC 9309). A crawler uses the group whose User-agent names its token, matched without regard to case. If several groups name it, their rules are combined. Only when no group names it does it fall back to User-agent: *. Among the rules that match a path, the longest one wins, and Allow wins a tie. An empty Disallow: blocks nothing.
Allowed, N paths restricted means the crawler can read your home page but some Disallow rules in its group cover other paths. That's often intended (admin, cart, search pages). The most common mistake it reveals is the opposite one: a site adds a dedicated AI group and forgets that the crawler now ignores every rule in the * group. The generator handles this for you by repeating your * rules in each allowed AI group.
Best practices for AI crawlers in robots.txt
Decide per purpose, not per vendor. If AI visibility matters to you, keep the search crawlers allowed even if you opt out of training. Avoid blanket "block all AI bots" settings in a CDN or firewall: they often don't separate search from training, and a blocked request never reaches your robots.txt. Anthropic also advises against blocking its bots by IP, because the bot then can't read your opt-out.
Each subdomain needs its own robots.txt. Vendors mention delays of up to about 24 hours before changes apply. Robots.txt is a request, not access control: OpenAI says its rules may not apply to ChatGPT-User, and Perplexity says Perplexity-User generally ignores them. User-agent strings can also be spoofed, so check server logs against the IP ranges vendors publish.
What robots.txt can't tell you
A clean robots.txt only means AI crawlers are allowed to read your pages. It doesn't tell you whether ChatGPT, Perplexity or Gemini actually mention your brand or cite your site when buyers ask about your category.
Citegram is a free Chrome extension that shows that outcome. It asks your prompts in ChatGPT, Gemini, Perplexity, Claude, DeepSeek, Grok and Google AI Overviews, using the accounts already signed in to your browser, then records whether your brand is mentioned, its position, and the sources each answer cited. Run your prompts before you change robots.txt, then again after the vendors' propagation window, to confirm you're still cited. The free plan tracks one prompt, and Pro adds unlimited prompts for $10 a month.
FAQ
How do I block GPTBot in robots.txt?
Add a group with User-agent: GPTBot followed by Disallow: /. OpenAI documents GPTBot as independent from OAI-SearchBot, so this opts you out of training without removing you from ChatGPT search. Keep OAI-SearchBot allowed if you want to stay visible there.
Will blocking AI crawlers hurt my Google rankings?
Blocking GPTBot, ClaudeBot or PerplexityBot doesn't affect Googlebot. Google also says Google-Extended doesn't affect inclusion in Google Search and isn't a ranking signal. What you can lose is visibility in AI assistants, if you block their search crawlers.
Which AI crawlers should I allow?
If you want to be cited by AI assistants, allow the search crawlers: OAI-SearchBot, Claude-SearchBot and PerplexityBot. Whether you allow training crawlers like GPTBot and ClaudeBot is a separate choice about model training. The "Allow AI search, block training" preset applies that split.
Do AI bots actually respect robots.txt?
Anthropic says its bots honor robots.txt, and OpenAI and Perplexity tell site owners to manage their search and training crawlers through it. User-triggered fetchers are the exception: OpenAI says robots.txt may not apply to ChatGPT-User, and Perplexity says Perplexity-User generally ignores it. See our full guide for the vendor sources.
Why is a bot still allowed when my User-agent: * group blocks everything?
It shouldn't be, unless a group names that bot. A crawler with its own group follows only that group and ignores User-agent: * entirely. The checker shows which group and line applied, so you can spot an Allow: / or an empty Disallow: that overrides your general rules.