Key takeaways
- Training crawlers (GPTBot, ClaudeBot, Google-Extended) and search crawlers (OAI-SearchBot, Claude-SearchBot, PerplexityBot) are separate, and vendors say you can allow one and block the other.
- OpenAI says sites opted out of OAI-SearchBot won't be shown in ChatGPT search answers, so blocking it has a visibility cost.
- Since 31 August 2026 Google has a Search Console control that excludes a site from AI Overviews and AI Mode. Google-Extended is a different control: it covers training and grounding and has no effect on Search.
- robots.txt is a request, not access control. User-triggered fetchers such as ChatGPT-User and Perplexity-User may not follow it, so enforcement needs a WAF, CDN or server rule.
- Check what any CDN preset does before you switch it on: Cloudflare said its training-block setting would also block multi-purpose crawlers such as Googlebot.
Block the crawlers that train AI models if you object to your content being used that way, but think twice before blocking the ones that power AI search. The two groups use different bots, so you can say no to one and yes to the other.
Most advice on this topic lumps every AI bot together. That’s how sites end up blocking a search bot by accident and then wondering why they’ve vanished from ChatGPT answers.
This guide takes each crawler’s behaviour from the company that runs it, checked against their documentation on 3 October 2026. Where a vendor hasn’t said something, we say so.
Which AI Crawlers Should You Know About?
The main AI crawlers come from OpenAI, Anthropic, Perplexity, Google, Apple, Amazon and Common Crawl, and each vendor documents a different purpose for each bot. The table sticks to what those vendors say publicly.
| Crawler | Operator | Type | Documented purpose |
|---|---|---|---|
| GPTBot | OpenAI | Training | Crawls content that may be used to train generative AI foundation models |
| OAI-SearchBot | OpenAI | Search | Surfaces websites in ChatGPT’s search features |
| ChatGPT-User | OpenAI | User-triggered | Visits pages for certain user actions in ChatGPT and Custom GPTs |
| ClaudeBot | Anthropic | Training | Collects web content that could contribute to model training |
| Claude-SearchBot | Anthropic | Search | Improves the quality of search results for users |
| Claude-User | Anthropic | User-triggered | Fetches pages when a Claude user asks a question |
| PerplexityBot | Perplexity | Search | Surfaces and links websites in Perplexity results; not used for foundation model training |
| Perplexity-User | Perplexity | User-triggered | May visit a page to answer a user’s question |
| Amazonbot | Amazon | Training | Improves Amazon products and services; may be used to train Amazon AI models |
| Amzn-SearchBot | Amazon | Search | Improves search experiences such as Alexa; doesn’t crawl for generative AI training |
| Amzn-User | Amazon | User-triggered | Fetches live information for requests such as Alexa questions |
| Google-Extended | Training control | Token that controls use of crawled content for Gemini training and grounding | |
| Applebot-Extended | Apple | Training control | Token that controls use of Applebot data to train Apple’s foundation models |
| CCBot | Common Crawl | Open dataset | Builds an open repository of web crawl data |
OpenAI: three main bots, three jobs
OpenAI’s crawler documentation lists OAI-SearchBot, GPTBot and ChatGPT-User, plus OAI-AdsBot, which only visits pages submitted as ads. It says each setting is independent: a webmaster can allow OAI-SearchBot to appear in search results while disallowing GPTBot.
If you allow both, OpenAI says it may use the results of one crawl for both purposes to avoid duplicate crawling. That’s a reason to be deliberate about GPTBot rather than leaving it to chance.
Anthropic: the same three-way split
Anthropic’s help article, dated 7 April 2026, names ClaudeBot, Claude-User and Claude-SearchBot. Restricting ClaudeBot signals that a site’s future materials should be excluded from training datasets.
Anthropic warns that disabling Claude-User may reduce your visibility for user-directed web search. For Claude-SearchBot it says disabling may reduce your visibility and accuracy in user search results.
Perplexity, Amazon and Common Crawl
Perplexity says PerplexityBot is designed to surface and link websites in search results and is not used to crawl content for AI foundation models. Perplexity-User supports user actions, and the documentation says it generally ignores robots.txt rules because a person requested the fetch.
Amazon documents three bots with independent settings. Amzn-SearchBot and Amzn-User don’t crawl for generative AI training, whereas Amazonbot may be used to train Amazon AI models.
Common Crawl describes itself as a non-profit that maintains an open repository of web crawl data. Its CCBot identifies itself as CCBot/2.0, and the project warns that other crawlers falsely identify themselves as CCBot.
Google and Apple: control tokens, not separate crawlers
Google-Extended has no HTTP user agent of its own: crawling uses existing Google user agents and the token only acts as a control in robots.txt. Google’s documentation says it manages whether crawled content can be used to train future Gemini models and for grounding, and that it does not impact inclusion in Google Search or act as a ranking signal.
Apple’s documentation says Applebot-Extended does not crawl webpages, and that pages which disallow it can still be included in search results. It only determines how data crawled by Applebot is used.
Which user-agent versions are documented?
Match on the bot name, not the version number, because OpenAI’s page says the version number may change. These are the versions each vendor listed on 3 October 2026.
| Bot | Documented user-agent token |
|---|---|
| GPTBot | GPTBot/1.4 |
| OAI-SearchBot | OAI-SearchBot/1.4 |
| ChatGPT-User | ChatGPT-User/1.0 |
| OAI-AdsBot | OAI-AdsBot/1.0 |
| PerplexityBot | PerplexityBot/1.0 |
| Perplexity-User | Perplexity-User/1.0 |
| Amazonbot | Amazonbot/0.1 |
| Amzn-SearchBot | Amzn-SearchBot/0.1 |
| Amzn-User | Amzn-User/0.1 |
| CCBot | CCBot/2.0 |
| ClaudeBot, Claude-SearchBot, Claude-User | No version published in Anthropic’s help article |
What’s the Difference Between Training, Search and User-Triggered Bots?
Training crawlers collect content to build or improve a model, search crawlers index pages so an AI product can find and cite them, and user-triggered fetchers visit a page only because someone asked. Cloudflare uses the same three-way split, which it calls training, search and agent.
Training crawlers
GPTBot, ClaudeBot, Amazonbot, Google-Extended and Applebot-Extended all sit here. Opting out tells the vendor your content shouldn’t be used to train its models.
Nothing in the vendor documents we read says a training opt-out changes where you appear in that vendor’s own search product. OpenAI says outright that its settings are independent.
Search and retrieval crawlers
OAI-SearchBot, Claude-SearchBot, PerplexityBot and Amzn-SearchBot index pages so an assistant can surface them. OpenAI says sites opted out of OAI-SearchBot won’t be shown in ChatGPT search answers, though they can still appear as navigational links.
User-triggered fetchers
ChatGPT-User, Claude-User, Perplexity-User and Amzn-User act only when a person asks. OpenAI says robots.txt rules may not apply to ChatGPT-User because the action is user-initiated, and that it isn’t used to decide whether content appears in Search.
Will Blocking AI Crawlers Cost You Citations and Traffic?
Blocking search crawlers can, and the vendors say so. Blocking training crawlers has no documented effect on citations, but nobody outside the AI labs can say what it does to how a model describes your brand.
What the vendors say about search bots
For OpenAI, the line is clear: opt out of OAI-SearchBot and you won’t be shown in ChatGPT search answers. For Anthropic, disabling Claude-SearchBot or Claude-User “may reduce” your visibility, and Perplexity recommends allowing PerplexityBot so your site appears in its results.
Amazon says that by permitting Amzn-SearchBot, your content becomes eligible to appear in search experiences such as Alexa. Notice the pattern: every search bot is described as the route into that product’s answers.
What nobody knows about training bots
A model that never saw your content may describe you less accurately, but no vendor publishes evidence either way. Treat any confident claim here, including ours, as opinion.
What we do know about retrieval is in our guide to how to get your brand cited by ChatGPT, which starts from the assumption that the search bot can reach you. An AI visibility audit is the quickest way to see how assistants describe you today.
What blocking does not cost you in Google Search
Blocking GPTBot, ClaudeBot or CCBot has nothing to do with Googlebot, because they are separate crawlers run by separate companies. Blocking Google-Extended doesn’t touch Search either, since Google says it has no effect on inclusion or ranking.
The risk is the opposite one: a careless User-agent: * rule or a CDN preset that catches Googlebot too. We come back to that in the enforcement section below.
Can You Opt Out of Google’s AI Overviews and AI Mode?
Yes: since 31 August 2026, Search Console has a Search generative AI control that excludes your site from AI Overviews, AI Mode and generative AI features in Discover. It doesn’t cover AI training, which is still the job of Google-Extended.
How the Search Console control works
Google’s Search generative AI control has three settings: include, exclude, or inherit from a parent property. Excluding your site means its links and content won’t appear in those features, and you won’t receive any traffic or impressions from them.
Google says the control isn’t used as a ranking or inclusion signal for other parts of Search, and that a change generally takes a few days to apply. It adds that content from other sites will still appear in those features, and it may look similar to yours.
Why Google built it
The UK’s Competition and Markets Authority imposed a conduct requirement on Google on 3 June 2026, which the CMA described as a world first: publishers can now opt out of their content powering AI features in Google search. Google was given nine months to implement all the changes.
Google’s own announcement on the same day said it was beginning to test the toggle with a subset of UK website owners, and an update confirms it reached all websites worldwide on 31 August 2026. Google’s AI features documentation, last updated in December 2025, predates all of this and points only to snippet controls and Google-Extended.
Which control does what?
Pick the control that matches the outcome you want, because they don’t overlap.
| What you want | Control | Effect on Google Search |
|---|---|---|
| Stop your content training future Gemini models | Google-Extended in robots.txt |
None: no effect on inclusion or ranking |
| Leave AI Overviews, AI Mode and Discover AI features | Search generative AI control in Search Console | No traffic or impressions from those features; not a ranking signal elsewhere |
| Limit the text Google shows from a page | nosnippet, data-nosnippet, max-snippet |
Changes what is shown in results |
| Disappear from Search completely | noindex |
Page is removed from Search |
Related: How Do Google AI Overviews Choose Their Sources?
How Do You Block AI Crawlers in robots.txt?
Add one User-agent group per bot with Disallow: /, and either leave the bots you want to keep out of the file or allow them explicitly. The syntax is the same for every vendor.
Block training bots only
This is the common middle path. It keeps you eligible for AI search while opting out of training.
User-agent: GPTBot
Disallow: /
User-agent: ClaudeBot
Disallow: /
User-agent: Amazonbot
Disallow: /
User-agent: Google-Extended
Disallow: /
User-agent: Applebot-Extended
Disallow: /
User-agent: CCBot
Disallow: /
Allow search bots explicitly
An explicit allow makes your intent obvious to the next person who edits the file. It also protects the search bots if someone later adds a blanket rule.
User-agent: OAI-SearchBot
Allow: /
User-agent: Claude-SearchBot
Allow: /
User-agent: PerplexityBot
Allow: /
User-agent: Amzn-SearchBot
Allow: /
RFC 9309 says a crawler obeys the group that matches its own name and falls back to the * group only if no group matches. So a named allow group beats a wildcard block for that bot.
Block part of a site
You can disallow a folder instead of the whole site, as in Apple’s own example of Disallow: /private/. That suits sites with a paywalled, licensed or members-only section.
User-agent: GPTBot
Disallow: /research/
Disallow: /members/
To see which AI engines can reach your site and what they say about your brand, book an AI visibility audit, from £500.
Which robots.txt mistakes cause the most trouble?
Most failures come from the details of the standard, not the AI bots. These come straight from RFC 9309, the IETF standard published in September 2022.
- Wrong host. The file applies to the host it sits on, so blog.example.com needs its own. Anthropic says to do this for every subdomain.
- Mixed-up case. Matching the user-agent name is case-insensitive, but matching paths should be case-sensitive, so
/Private/and/private/are different. - Rule conflicts. The most specific (longest) matching rule wins, and if an allow and a disallow are equivalent, the allow wins.
- Server errors. A 4xx response lets crawlers access anything, while a 5xx or unreachable server makes them assume a complete disallow.
- Slow changes. Crawlers shouldn’t use a cached copy for more than 24 hours, and OpenAI, Perplexity and Amazon all quote about a day for changes to apply.
Does robots.txt Stop AI Crawlers on Its Own?
Not by itself: compliant bots honour it, but the standard says its rules are not a form of access authorization, and some fetchers and rogue crawlers ignore it. If a bot matters enough to block, enforce the block somewhere stronger.
What the standard says
RFC 9309 states that the Robots Exclusion Protocol “is not a substitute for valid content security measures”. It also warns that listing paths in robots.txt exposes them publicly, so a Disallow line can point curious visitors straight at the thing you wanted hidden.
Anthropic adds a related warning: blocking its IP addresses may not guarantee an opt-out, because doing so stops the bot reading your robots.txt in the first place. Keep robots.txt readable and add other measures alongside it.
Which fetchers may ignore it
OpenAI says robots.txt rules may not apply to ChatGPT-User because users initiate the action. Perplexity says Perplexity-User generally ignores robots.txt rules, and Amazon says Amzn-User may not follow all directives.
There’s also the question of bots that don’t identify themselves honestly. Cloudflare reported in August 2025 that it saw Perplexity changing its user agent and network when blocked, and removed Perplexity as a verified bot; that is Cloudflare’s account of its own tests, not a finding we could confirm independently.
How to enforce a block
For real enforcement, block at the WAF, CDN or server. Helen Pollitt, Head of SEO at Getty Images, wrote in Search Engine Journal that the higher up the stack you go the better: WAF if you can, CDN if you can’t, server as a last resort.
Cloudflare’s AI Crawl Control lets you set allow or block rules for individual crawlers and track which ones violate your robots.txt. Perplexity’s documentation recommends combining user-agent matching with its published IP ranges if you want to allow its bots through a firewall.
What should you check before using a CDN preset?
Presets can catch more than you intend. Cloudflare said in a 1 July 2026 post that, from 15 September 2026, multi-purpose crawlers that combine search with training would be treated according to all their behaviours, so Googlebot, Applebot and Bingbot would be blocked for customers who had chosen to block training.
We haven’t tested how that plays out on live sites. If you use Cloudflare, open the AI settings and confirm that the search crawlers you rely on are still allowed before you switch on any training block.
Related: What Is llms.txt and Does Your Site Need One?
How Do You Check Which AI Bots Visit Your Site?
Search your server or CDN logs for the bot names, then check any IP address against the vendor’s published list, because a user-agent string can be faked. Ten minutes in the logs tells you more than any list of bots.
What to look for in the logs
Count hits per bot name, then look at the response codes. A bot you’ve blocked in robots.txt that still returns 200 on content pages is worth investigating.
This command counts hits per AI bot in a standard access log:
grep -oE "GPTBot|OAI-SearchBot|ChatGPT-User|ClaudeBot|Claude-SearchBot|Claude-User|PerplexityBot|Perplexity-User|Amzn-SearchBot|Amzn-User|Amazonbot|CCBot" access.log | sort | uniq -c | sort -rn
Applebot-Extended and Google-Extended won’t appear, because neither has a user agent of its own. Apple says Applebot-Extended doesn’t crawl, and Google says Google-Extended crawls with existing user agents.
How to verify a bot is genuine
OpenAI, Perplexity, Common Crawl and Apple publish IP lists or reverse DNS checks, and Anthropic’s help article links to its own IP list. Use them, because spoofing a user-agent string costs nothing.
| Bot | How to verify |
|---|---|
| OAI-SearchBot, GPTBot, ChatGPT-User | IP lists at openai.com/searchbot.json, gptbot.json and chatgpt-user.json |
| Claude bots | IP list at claude.com/crawling/bots.json |
| PerplexityBot, Perplexity-User | IP lists at perplexity.com/perplexitybot.json and perplexity-user.json |
| CCBot | Reverse DNS ending crawl.commoncrawl.org, or the list at index.commoncrawl.org/ccbot.json |
| Applebot | Reverse DNS in the *.applebot.apple.com domain, or Apple’s published CIDR file |
A worked example
Imagine a log line like this one on a small business site:
203.0.113.7 - - [03/Oct/2026:09:14:22 +0000] "GET /guides/ HTTP/1.1" 200 18452 "-" "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.4; +https://openai.com/gptbot"
The user agent says GPTBot, but if 203.0.113.7 isn’t in OpenAI’s gptbot.json list, you’re looking at something pretending to be GPTBot. The 200 response would also tell you any robots.txt rule is being ignored, which is the cue to add a firewall rule rather than another robots.txt line.
Related: How Do You Measure AI Search Visibility?
Should You Block AI Crawlers or Allow Them?
For most businesses that want leads, allow the search and retrieval bots and make a deliberate choice about the training bots. Publishers whose content is the product have a stronger case for blocking more.
| Type of site | Training bots | Search and retrieval bots | Reason |
|---|---|---|---|
| Local service business or agency | Your call | Allow | Being named in an answer is the whole point |
| Lead generation or comparison site | Your call | Allow | Visitors arrive from answers, so blocking trades visits for little |
| Paywalled research or licensed data | Block | Allow free pages only | The content is the asset |
| Original news or reporting | Block, and consider a firewall rule | Allow selectively | Licensing power matters more than reach |
Allow if you rely on being found
Picture a Cheltenham accountancy firm whose best leads come from people asking an assistant for a local specialist. Blocking OAI-SearchBot or PerplexityBot there trades visibility for very little.
The same logic runs through generative engine optimisation, which assumes the bots can get in.
Block if your content is the asset
If you sell research behind a paywall, opting out of training is a reasonable position. Do it knowing it’s a request, and back it with firewall rules if the content really is sensitive.
Review on a schedule
Bot names, versions and policies change, and so do Google’s controls, as the Search Console toggle shows. Diarise a review each quarter and compare your robots.txt against the vendor pages.
Crawler access is the first thing to rule out in any AI search or SEO project, because nothing else works if the bots can’t get in.
If you’d like an outside view of what AI engines can reach and say about your brand, get an AI visibility audit from £500.
| Training crawlers | Search and retrieval bots | |
|---|---|---|
| Examples | GPTBot, ClaudeBot, Amazonbot | OAI-SearchBot, Claude-SearchBot, PerplexityBot |
| What they feed | Model training | Search indexes and cited answers |
| If you block them | Content opted out of training | You can drop out of AI answers |
| Typical decision | A content and licensing choice | A visibility choice |
FAQs
Will blocking AI crawlers hurt my Google rankings?
Not if you only block the AI vendors' bots. Google says Google-Extended doesn't impact a site's inclusion in Google Search nor is it used as a ranking signal, and OpenAI, Anthropic and Perplexity run separate bots from Googlebot. Blocking Googlebot itself would remove you from Search, so never do that by accident.
How long does a robots.txt change take to apply?
OpenAI and Amazon both say about 24 hours, and Perplexity says up to 24 hours. Anthropic doesn't publish a figure. The robots.txt standard itself says crawlers shouldn't use a cached copy for more than 24 hours unless the file is unreachable.
Do I need to block each subdomain separately?
Yes. Anthropic says to apply the rule for every subdomain you want to opt out, and Amazon says its bots read robots.txt at the host level and honour the rules under each host. Treat blog.example.com and www.example.com as two separate files.
What happens if my robots.txt returns an error?
It depends on the error. Under RFC 9309, a 4xx response means the crawler may access anything, while a 5xx response or an unreachable server means it must assume a complete disallow. A broken server can therefore hide your whole site from well-behaved bots.
Can I slow AI crawlers down instead of blocking them?
Only with some of them. Anthropic says its bots support the non-standard Crawl-delay directive, while Apple and Amazon both say their crawlers don't follow it. Neither OpenAI nor Perplexity mention it in their crawler documentation.
Should I block CCBot?
CCBot belongs to Common Crawl, a non-profit that maintains an open repository of web crawl data that anyone can analyse. If you don't want to be in that repository, Common Crawl's own documentation gives a Disallow rule for CCBot, and we found no statement from Common Crawl about what happens to pages already archived.
What is Cloudflare's pay per crawl?
It's a Cloudflare feature that lets AI crawlers pay to access your content. Cloudflare's documentation lists it as a private beta, so most sites can't use it yet.
Sources
- Overview of OpenAI Crawlers, OpenAI
- Does Anthropic crawl data from the web, and how can site owners block the crawler?, Anthropic Help Center
- Perplexity Crawlers, Perplexity
- Google's common crawlers, Google Search Central
- Search generative AI control, Search Console Help
- New opportunities, control and insights for website owners, Google, The Keyword
- AI features and your website, Google Search Central
- CMA secures fairer deal for publishers and improves Google search services in UK, GOV.UK, Competition and Markets Authority
- Google search publisher conduct requirement, GOV.UK, Competition and Markets Authority
- About Applebot, Apple Support
- About Amazonbot, Amazon
- CCBot, Common Crawl
- RFC 9309: Robots Exclusion Protocol, IETF
- Your site, your rules: new AI traffic options for all customers, Cloudflare
- Perplexity is using stealth, undeclared crawlers to evade website no-crawl directives, Cloudflare
- AI Crawl Control, Cloudflare Docs
- Should I Block AI Crawlers At Robots.txt Or Server Level? (Ask An SEO), Search Engine Journal, Helen Pollitt
Related services