Allow or block well-known AI crawlers. robots.txt is a voluntary standard: well-behaved crawlers follow it, others may not. Blocking Google-Extended does not remove you from Google Search.
| Crawler | Operator and purpose | Rule |
|---|---|---|
| GPTBot | OpenAI · training data | |
| OAI-SearchBot | OpenAI · ChatGPT search results | |
| ChatGPT-User | OpenAI · fetches on a user’s request | |
| ClaudeBot | Anthropic · training data | |
| Claude-SearchBot | Anthropic · search quality | |
| Claude-User | Anthropic · fetches on a user’s request | |
| PerplexityBot | Perplexity · search index | |
| Perplexity-User | Perplexity · fetches on a user’s request | |
| Google-Extended | Google · Gemini training and grounding; does not affect Search | |
| Applebot-Extended | Apple · Apple AI training | |
| CCBot | Common Crawl · open web dataset | |
| Meta-ExternalAgent | Meta · Meta AI training |
Review your existing robots.txt first and merge rather than overwrite. Crawler names and behaviour change, so check each operator’s documentation.