Technical
AI Crawlers: The 2026 List of AI Bot User Agents
A verified list of AI crawler user agents (GPTBot, ClaudeBot, PerplexityBot and more), what each does, plus robots.txt rules to allow search, block training.
- Published
- Last updated
- Reading time
- 9 min read
AI crawlers are the bots AI companies use to fetch web pages, either to collect training data for their models, to build a search index their assistant can cite, or to open a page live because a user asked. The practical rule: allow the search and user-triggered bots if you want to be cited in AI answers, and decide separately whether to allow the training bots. Each company uses different user agents for these jobs, so a single robots.txt rule rarely does what people expect.
Below is a reference list of the major AI crawler user agents in 2026, verified against each vendor's own documentation on September 25, 2026, followed by robots.txt examples and a way to check your setup.
What are the three types of AI crawlers?
An AI web crawler is an automated client that fetches pages for an AI product. They fall into three groups, and the group determines what blocking it costs you.
- Training crawlers collect content that may be used to train future models. Examples: GPTBot, ClaudeBot, Meta-ExternalAgent, CCBot. Blocking them keeps future content out of training sets. It does not remove you from AI search results.
- Search crawlers build an index the assistant retrieves from when it answers with web results. Examples: OAI-SearchBot, Claude-SearchBot, PerplexityBot, Amzn-SearchBot, and Googlebot for Google's AI features. Blocking them removes your pages from those answers and citations.
- User-triggered fetchers open a specific page because a person asked the assistant to, or because the assistant needed it mid-answer. Examples: ChatGPT-User, Claude-User, Perplexity-User, Meta-ExternalFetcher. Several vendors say these may not follow robots.txt, because a human initiated the request.
There is also a fourth, odd category: control tokens that do not crawl at all. Google-Extended and Applebot-Extended are names you use in robots.txt to opt out of AI training, while the normal Googlebot and Applebot do the fetching.
The AI crawler user agent list
| User agent token | Company | Type | Follows robots.txt | Blocking it means |
|---|---|---|---|---|
| GPTBot | OpenAI | Training | Yes | Content excluded from model training |
| OAI-SearchBot | OpenAI | Search | Yes | Not surfaced in ChatGPT search results |
| ChatGPT-User | OpenAI | User-triggered | May not apply | Limited effect via robots.txt |
| OAI-AdsBot | OpenAI | Ad checks | Only visits pages submitted as ads | Not relevant unless you run ChatGPT ads |
| ClaudeBot | Anthropic | Training | Yes | Future content excluded from training |
| Claude-SearchBot | Anthropic | Search | Yes | Not indexed for Claude's search |
| Claude-User | Anthropic | User-triggered | Yes (per Anthropic) | Claude cannot fetch your page for users |
| PerplexityBot | Perplexity | Search | Yes | Not surfaced or linked in Perplexity |
| Perplexity-User | Perplexity | User-triggered | Generally ignores | Needs firewall rules to block |
| Googlebot | Search (incl. AI Overviews, AI Mode) | Yes | Removed from Google Search and its AI features | |
| Google-Extended | Control token (training, grounding) | Yes | Not used for Gemini training or grounding; no effect on Search | |
| Applebot | Apple | Search (Spotlight, Siri, Safari) | Yes | Not in Apple search features |
| Applebot-Extended | Apple | Control token (training) | Yes | Not used to train Apple foundation models |
| Amazonbot | Amazon | Products, may train models | Yes | Not used by Amazon, including model training |
| Amzn-SearchBot | Amazon | Search (e.g. Alexa) | Yes | Not in Amazon search experiences |
| Amzn-User | Amazon | User-triggered | May not follow all rules | Limited effect via robots.txt |
| Meta-ExternalAgent | Meta | Training and indexing | Yes | Excluded from Meta's crawl |
| Meta-ExternalFetcher | Meta | User-triggered | May bypass | Limited effect via robots.txt |
| CCBot | Common Crawl | Open web archive | Yes | Excluded from Common Crawl's public data set |
| Bytespider | ByteDance | Undocumented | No official statement | See note below |
OpenAI: GPTBot, OAI-SearchBot, ChatGPT-User
OpenAI's crawler documentation lists four agents. GPTBot crawls content that may be used to make its generative AI foundation models more useful and safe. OAI-SearchBot is used to surface websites in ChatGPT's search features. ChatGPT-User handles certain user actions in ChatGPT and Custom GPTs; OpenAI notes that because these actions are initiated by a user, robots.txt rules may not apply. OAI-AdsBot only visits pages submitted as ads, and OpenAI says that data is not used for training.
OpenAI says robots.txt changes for OAI-SearchBot take about 24 hours to take effect. It publishes IP ranges for GPTBot, OAI-SearchBot and ChatGPT-User, which you should use to verify traffic claiming to be these bots.
Anthropic: ClaudeBot, Claude-SearchBot, Claude-User
Anthropic's help article describes three bots. ClaudeBot collects web content that could contribute to model training; disallowing it signals that future content should be excluded from training data. Claude-SearchBot crawls to improve search result quality for Claude users. Claude-User fetches pages when a person asks Claude a question. Anthropic warns that blocking Claude-SearchBot or Claude-User may reduce your site's visibility in Claude's answers.
Anthropic supports the non-standard Crawl-delay directive for rate limiting and publishes its crawler IP addresses.
Perplexity: PerplexityBot, Perplexity-User
Perplexity's bot documentation says PerplexityBot is designed to surface and link websites in Perplexity's search results and is not used to crawl content for AI foundation models. Perplexity-User supports user actions: when someone asks Perplexity a question it may visit a page to answer, and Perplexity says this fetcher generally ignores robots.txt because a user initiated the request. Robots.txt changes can take up to 24 hours. IP lists: PerplexityBot, Perplexity-User.
Google: Googlebot and Google-Extended
This is where most robots.txt files get it wrong. Google's AI Overviews and AI Mode do not use a separate AI crawler. Google's AI features documentation says a page only needs to be indexed and eligible to show with a snippet in Google Search, and that the usual controls apply: nosnippet, data-nosnippet, max-snippet and noindex.
Google-Extended is a separate product token. Per Google's crawler list, it controls whether content Google crawls may be used to train future Gemini models and for grounding in Gemini Apps and the Vertex AI API for Gemini. Google states it does not affect inclusion in Google Search and is not a ranking signal. It has no user agent string of its own: crawling is done by existing Google user agents, so you will never see "Google-Extended" in your logs.
So blocking Google-Extended does not keep you out of AI Overviews, and blocking Googlebot to avoid AI Overviews removes you from Google Search entirely.
Apple: Applebot and Applebot-Extended
Apple's documentation says Applebot powers search features in Spotlight, Siri and Safari, and that crawled data may also be used to train Apple foundation models behind Apple Intelligence. Applebot-Extended does not crawl; disallowing it opts your content out of that training, and Apple says pages that disallow it can still appear in search results. One detail: if your robots.txt does not mention Applebot but does mention Googlebot, Applebot follows the Googlebot rules.
Amazon, Meta and Common Crawl
Amazon runs Amazonbot (improves products and may be used to train Amazon AI models), Amzn-SearchBot (search experiences such as Alexa; not used for training) and Amzn-User (user actions; may not follow all robots.txt directives). If your robots.txt does not mention Amzn-SearchBot but allows other search bots, Amazon says it follows the rules given to those bots.
Meta runs Meta-ExternalAgent for uses such as training foundation models or indexing content, and Meta-ExternalFetcher for user-requested fetches, which may bypass robots.txt. The separate facebookexternalhit bot builds link previews when a URL is shared.
Common Crawl's CCBot builds an open repository of web crawl data anyone can use, which makes it an indirect route into many AI data sets. It respects robots.txt and publishes its IP ranges for verification.
Bytespider
Bytespider is ByteDance's crawler. We could not find official documentation, an IP list or a robots.txt statement from ByteDance, so we have nothing authoritative to link. If it causes load problems, block it at the firewall or CDN by user agent rather than relying on robots.txt alone.
Microsoft Copilot
Copilot's web answers are tied to Bing: Bing's webmaster guidelines cover how content is surfaced across Bing search and Copilot, so the crawler that matters is bingbot, controlled through your normal robots.txt rules. Bing Webmaster Tools added an AI Performance report in February 2026 that shows citations across Microsoft Copilot and AI-generated summaries in Bing.
How to write robots.txt rules for AI crawlers
Two rules of robots.txt matter here, both from Google's robots.txt specification, which major crawlers follow:
- A crawler obeys only the most specific group that matches its name. If you create a
User-agent: GPTBotgroup, GPTBot ignores everything underUser-agent: *. Repeat any shared disallows inside each named group. - The file only applies to the host and protocol it sits on.
blog.example.comneeds its own robots.txt.
Allow AI search, block AI training
This is the most common setup for brands that want to be cited but do not want future content used for training:
# Training crawlers and training control tokens: blocked
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: Meta-ExternalAgent
User-agent: CCBot
User-agent: Amazonbot
Disallow: /
# AI search and user-triggered bots: allowed
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: PerplexityBot
User-agent: Perplexity-User
User-agent: Amzn-SearchBot
Allow: /
Disallow: /admin/
Disallow: /cart/
# Everyone else, including Googlebot and Bingbot
User-agent: *
Disallow: /admin/
Disallow: /cart/
Sitemap: https://www.example.com/sitemap.xml
Note the trade-off with Google-Extended: Google says it also covers grounding in Gemini Apps and the Vertex AI API, so blocking it may reduce how Gemini uses your content beyond training. Decide based on how much Gemini traffic matters to you.
Allow everything
If you want maximum AI exposure, you do not need AI-specific lines at all. A permissive User-agent: * group lets every compliant bot in. Adding explicit Allow groups for search bots does no harm and documents intent.
Block everything AI
User-agent: GPTBot
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: ClaudeBot
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: PerplexityBot
User-agent: Perplexity-User
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: Meta-ExternalAgent
User-agent: Meta-ExternalFetcher
User-agent: Amazonbot
User-agent: CCBot
User-agent: Bytespider
Disallow: /
Expect this to remove you from ChatGPT, Claude and Perplexity answers, and remember that user-triggered fetchers may still visit. It does not affect Google AI Overviews, which follow Googlebot. The robots.txt generator builds any of these variants with the current token list.
Why robots.txt is not the whole story
Firewalls and CDNs. Many sites block AI bots at the CDN or bot-management layer without anyone deciding to. Your robots.txt can say "allow" while a firewall returns 403. Check both.
User-triggered fetchers. OpenAI, Perplexity, Amazon and Meta all say their user-triggered agents may not follow robots.txt. If you truly need to stop them, use firewall rules matched on user agent and verified IP ranges.
Spoofed user agents. Anyone can send "GPTBot" in a header. Verify against the published IP lists above before trusting log data or allowlisting. Perplexity recommends combining user agent matching with IP verification and refreshing IP lists automatically.
JavaScript rendering. Vercel's December 2024 analysis found that none of the major AI crawlers it observed, including GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot and PerplexityBot, rendered JavaScript. If your main content loads client-side, these bots may see an empty page even when access is allowed. Use server-side rendering or static generation for pages you want cited, and check what bots receive with the crawler view tool.
llms.txt is separate. An llms.txt file is a Markdown index of your key pages for language models, proposed at llmstxt.org. It does not control access; robots.txt does. See what is llms.txt for when it is worth adding.
How to check your AI crawler access
An AI bot checker tests whether each AI user agent can reach your pages. Do it in this order:
- Run the AI crawler checker on your domain. It reads your robots.txt and reports allowed or blocked for each AI user agent listed above.
- Request a page as each bot. A curl request with the bot's user agent string (for example, the GPTBot string from OpenAI's docs) shows whether your CDN returns 200 or 403. This catches firewall blocks robots.txt tests miss.
- Read your server logs. Filter by the tokens in the table. No OAI-SearchBot or PerplexityBot hits over several weeks on a public site usually means something upstream blocks them.
- Check rendered HTML. Fetch a key page without JavaScript and confirm the main text, headings and prices are present.
- Re-check after changes. OpenAI and Perplexity both say robots.txt updates can take about 24 hours to apply.
Once bots can reach your pages, the next question is whether AI engines actually mention you. The AI visibility checker runs a baseline, and how to optimize content for AI search covers what to change on the pages themselves.
Frequently asked questions
What is GPTBot?
GPTBot is OpenAI's crawler for collecting content that may be used to train its foundation models. It respects robots.txt. Blocking GPTBot does not remove you from ChatGPT search results; that is controlled by a separate bot, OAI-SearchBot.
If I block GPTBot, will my site disappear from ChatGPT?
No. OpenAI uses OAI-SearchBot to surface websites in ChatGPT's search features and ChatGPT-User for pages a user asks ChatGPT to open. You can disallow GPTBot for training and still allow OAI-SearchBot so your pages can be cited in ChatGPT search answers.
Does Google-Extended stop my site appearing in AI Overviews?
No. Google says Google-Extended does not affect inclusion in Google Search and is not a ranking signal. AI Overviews and AI Mode use the normal Googlebot index. To limit how a page is shown in AI Overviews you use snippet controls such as nosnippet, data-nosnippet or max-snippet, or noindex.
Do AI crawlers execute JavaScript?
Generally not. Vercel's December 2024 analysis of its network found that none of the major AI crawlers it measured, including GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot and PerplexityBot, rendered JavaScript. Content that only appears after client-side rendering may be invisible to them, so serve important text in the initial HTML.
How do I check which AI bots can access my site?
Read your robots.txt for each user agent token, then check your CDN or firewall rules, since many sites block bots there without realising it. The aeotime AI crawler checker tests the common AI user agents against your robots.txt in one run. For real traffic, filter your server logs by user agent and verify IPs against each vendor's published list.