Skip to content

AI search

Should You Block AI Crawlers Like GPTBot?

Key takeaways

  • Training crawlers (GPTBot, ClaudeBot, Google-Extended) and search crawlers (OAI-SearchBot, Claude-SearchBot, PerplexityBot) are separate, and vendors say you can allow one and block the other.
  • OpenAI says sites opted out of OAI-SearchBot won't be shown in ChatGPT search answers, so blocking it has a visibility cost.
  • Since 31 August 2026 Google has a Search Console control that excludes a site from AI Overviews and AI Mode. Google-Extended is a different control: it covers training and grounding and has no effect on Search.
  • robots.txt is a request, not access control. User-triggered fetchers such as ChatGPT-User and Perplexity-User may not follow it, so enforcement needs a WAF, CDN or server rule.
  • Check what any CDN preset does before you switch it on: Cloudflare said its training-block setting would also block multi-purpose crawlers such as Googlebot.

Block the crawlers that train AI models if you object to your content being used that way, but think twice before blocking the ones that power AI search. The two groups use different bots, so you can say no to one and yes to the other.

Most advice on this topic lumps every AI bot together. That’s how sites end up blocking a search bot by accident and then wondering why they’ve vanished from ChatGPT answers.

This guide takes each crawler’s behaviour from the company that runs it, checked against their documentation on 3 October 2026. Where a vendor hasn’t said something, we say so.

Which AI Crawlers Should You Know About?

The main AI crawlers come from OpenAI, Anthropic, Perplexity, Google, Apple, Amazon and Common Crawl, and each vendor documents a different purpose for each bot. The table sticks to what those vendors say publicly.

Crawler Operator Type Documented purpose
GPTBot OpenAI Training Crawls content that may be used to train generative AI foundation models
OAI-SearchBot OpenAI Search Surfaces websites in ChatGPT’s search features
ChatGPT-User OpenAI User-triggered Visits pages for certain user actions in ChatGPT and Custom GPTs
ClaudeBot Anthropic Training Collects web content that could contribute to model training
Claude-SearchBot Anthropic Search Improves the quality of search results for users
Claude-User Anthropic User-triggered Fetches pages when a Claude user asks a question
PerplexityBot Perplexity Search Surfaces and links websites in Perplexity results; not used for foundation model training
Perplexity-User Perplexity User-triggered May visit a page to answer a user’s question
Amazonbot Amazon Training Improves Amazon products and services; may be used to train Amazon AI models
Amzn-SearchBot Amazon Search Improves search experiences such as Alexa; doesn’t crawl for generative AI training
Amzn-User Amazon User-triggered Fetches live information for requests such as Alexa questions
Google-Extended Google Training control Token that controls use of crawled content for Gemini training and grounding
Applebot-Extended Apple Training control Token that controls use of Applebot data to train Apple’s foundation models
CCBot Common Crawl Open dataset Builds an open repository of web crawl data

OpenAI: three main bots, three jobs

OpenAI’s crawler documentation lists OAI-SearchBot, GPTBot and ChatGPT-User, plus OAI-AdsBot, which only visits pages submitted as ads. It says each setting is independent: a webmaster can allow OAI-SearchBot to appear in search results while disallowing GPTBot.

If you allow both, OpenAI says it may use the results of one crawl for both purposes to avoid duplicate crawling. That’s a reason to be deliberate about GPTBot rather than leaving it to chance.

Anthropic: the same three-way split

Anthropic’s help article, dated 7 April 2026, names ClaudeBot, Claude-User and Claude-SearchBot. Restricting ClaudeBot signals that a site’s future materials should be excluded from training datasets.

Anthropic warns that disabling Claude-User may reduce your visibility for user-directed web search. For Claude-SearchBot it says disabling may reduce your visibility and accuracy in user search results.

Perplexity, Amazon and Common Crawl

Perplexity says PerplexityBot is designed to surface and link websites in search results and is not used to crawl content for AI foundation models. Perplexity-User supports user actions, and the documentation says it generally ignores robots.txt rules because a person requested the fetch.

Amazon documents three bots with independent settings. Amzn-SearchBot and Amzn-User don’t crawl for generative AI training, whereas Amazonbot may be used to train Amazon AI models.

Common Crawl describes itself as a non-profit that maintains an open repository of web crawl data. Its CCBot identifies itself as CCBot/2.0, and the project warns that other crawlers falsely identify themselves as CCBot.

Google and Apple: control tokens, not separate crawlers

Google-Extended has no HTTP user agent of its own: crawling uses existing Google user agents and the token only acts as a control in robots.txt. Google’s documentation says it manages whether crawled content can be used to train future Gemini models and for grounding, and that it does not impact inclusion in Google Search or act as a ranking signal.

Apple’s documentation says Applebot-Extended does not crawl webpages, and that pages which disallow it can still be included in search results. It only determines how data crawled by Applebot is used.

Which user-agent versions are documented?

Match on the bot name, not the version number, because OpenAI’s page says the version number may change. These are the versions each vendor listed on 3 October 2026.

Bot Documented user-agent token
GPTBot GPTBot/1.4
OAI-SearchBot OAI-SearchBot/1.4
ChatGPT-User ChatGPT-User/1.0
OAI-AdsBot OAI-AdsBot/1.0
PerplexityBot PerplexityBot/1.0
Perplexity-User Perplexity-User/1.0
Amazonbot Amazonbot/0.1
Amzn-SearchBot Amzn-SearchBot/0.1
Amzn-User Amzn-User/0.1
CCBot CCBot/2.0
ClaudeBot, Claude-SearchBot, Claude-User No version published in Anthropic’s help article

What’s the Difference Between Training, Search and User-Triggered Bots?

Training crawlers collect content to build or improve a model, search crawlers index pages so an AI product can find and cite them, and user-triggered fetchers visit a page only because someone asked. Cloudflare uses the same three-way split, which it calls training, search and agent.

Training crawlers

GPTBot, ClaudeBot, Amazonbot, Google-Extended and Applebot-Extended all sit here. Opting out tells the vendor your content shouldn’t be used to train its models.

Nothing in the vendor documents we read says a training opt-out changes where you appear in that vendor’s own search product. OpenAI says outright that its settings are independent.

Search and retrieval crawlers

OAI-SearchBot, Claude-SearchBot, PerplexityBot and Amzn-SearchBot index pages so an assistant can surface them. OpenAI says sites opted out of OAI-SearchBot won’t be shown in ChatGPT search answers, though they can still appear as navigational links.

User-triggered fetchers

ChatGPT-User, Claude-User, Perplexity-User and Amzn-User act only when a person asks. OpenAI says robots.txt rules may not apply to ChatGPT-User because the action is user-initiated, and that it isn’t used to decide whether content appears in Search.

Will Blocking AI Crawlers Cost You Citations and Traffic?

Blocking search crawlers can, and the vendors say so. Blocking training crawlers has no documented effect on citations, but nobody outside the AI labs can say what it does to how a model describes your brand.

What the vendors say about search bots

For OpenAI, the line is clear: opt out of OAI-SearchBot and you won’t be shown in ChatGPT search answers. For Anthropic, disabling Claude-SearchBot or Claude-User “may reduce” your visibility, and Perplexity recommends allowing PerplexityBot so your site appears in its results.

Amazon says that by permitting Amzn-SearchBot, your content becomes eligible to appear in search experiences such as Alexa. Notice the pattern: every search bot is described as the route into that product’s answers.

What nobody knows about training bots

A model that never saw your content may describe you less accurately, but no vendor publishes evidence either way. Treat any confident claim here, including ours, as opinion.

What we do know about retrieval is in our guide to how to get your brand cited by ChatGPT, which starts from the assumption that the search bot can reach you. An AI visibility audit is the quickest way to see how assistants describe you today.

Blocking GPTBot, ClaudeBot or CCBot has nothing to do with Googlebot, because they are separate crawlers run by separate companies. Blocking Google-Extended doesn’t touch Search either, since Google says it has no effect on inclusion or ranking.

The risk is the opposite one: a careless User-agent: * rule or a CDN preset that catches Googlebot too. We come back to that in the enforcement section below.

Can You Opt Out of Google’s AI Overviews and AI Mode?

Yes: since 31 August 2026, Search Console has a Search generative AI control that excludes your site from AI Overviews, AI Mode and generative AI features in Discover. It doesn’t cover AI training, which is still the job of Google-Extended.

How the Search Console control works

Google’s Search generative AI control has three settings: include, exclude, or inherit from a parent property. Excluding your site means its links and content won’t appear in those features, and you won’t receive any traffic or impressions from them.

Google says the control isn’t used as a ranking or inclusion signal for other parts of Search, and that a change generally takes a few days to apply. It adds that content from other sites will still appear in those features, and it may look similar to yours.

Why Google built it

The UK’s Competition and Markets Authority imposed a conduct requirement on Google on 3 June 2026, which the CMA described as a world first: publishers can now opt out of their content powering AI features in Google search. Google was given nine months to implement all the changes.

Google’s own announcement on the same day said it was beginning to test the toggle with a subset of UK website owners, and an update confirms it reached all websites worldwide on 31 August 2026. Google’s AI features documentation, last updated in December 2025, predates all of this and points only to snippet controls and Google-Extended.

Which control does what?

Pick the control that matches the outcome you want, because they don’t overlap.

What you want Control Effect on Google Search
Stop your content training future Gemini models Google-Extended in robots.txt None: no effect on inclusion or ranking
Leave AI Overviews, AI Mode and Discover AI features Search generative AI control in Search Console No traffic or impressions from those features; not a ranking signal elsewhere
Limit the text Google shows from a page nosnippet, data-nosnippet, max-snippet Changes what is shown in results
Disappear from Search completely noindex Page is removed from Search

Related: How Do Google AI Overviews Choose Their Sources?

How Do You Block AI Crawlers in robots.txt?

Add one User-agent group per bot with Disallow: /, and either leave the bots you want to keep out of the file or allow them explicitly. The syntax is the same for every vendor.

Block training bots only

This is the common middle path. It keeps you eligible for AI search while opting out of training.

User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: Amazonbot
Disallow: /

User-agent: Google-Extended
Disallow: /

User-agent: Applebot-Extended
Disallow: /

User-agent: CCBot
Disallow: /

Allow search bots explicitly

An explicit allow makes your intent obvious to the next person who edits the file. It also protects the search bots if someone later adds a blanket rule.

User-agent: OAI-SearchBot
Allow: /

User-agent: Claude-SearchBot
Allow: /

User-agent: PerplexityBot
Allow: /

User-agent: Amzn-SearchBot
Allow: /

RFC 9309 says a crawler obeys the group that matches its own name and falls back to the * group only if no group matches. So a named allow group beats a wildcard block for that bot.

Block part of a site

You can disallow a folder instead of the whole site, as in Apple’s own example of Disallow: /private/. That suits sites with a paywalled, licensed or members-only section.

User-agent: GPTBot
Disallow: /research/
Disallow: /members/

To see which AI engines can reach your site and what they say about your brand, book an AI visibility audit, from £500.

Which robots.txt mistakes cause the most trouble?

Most failures come from the details of the standard, not the AI bots. These come straight from RFC 9309, the IETF standard published in September 2022.

  • Wrong host. The file applies to the host it sits on, so blog.example.com needs its own. Anthropic says to do this for every subdomain.
  • Mixed-up case. Matching the user-agent name is case-insensitive, but matching paths should be case-sensitive, so /Private/ and /private/ are different.
  • Rule conflicts. The most specific (longest) matching rule wins, and if an allow and a disallow are equivalent, the allow wins.
  • Server errors. A 4xx response lets crawlers access anything, while a 5xx or unreachable server makes them assume a complete disallow.
  • Slow changes. Crawlers shouldn’t use a cached copy for more than 24 hours, and OpenAI, Perplexity and Amazon all quote about a day for changes to apply.

Does robots.txt Stop AI Crawlers on Its Own?

Not by itself: compliant bots honour it, but the standard says its rules are not a form of access authorization, and some fetchers and rogue crawlers ignore it. If a bot matters enough to block, enforce the block somewhere stronger.

What the standard says

RFC 9309 states that the Robots Exclusion Protocol “is not a substitute for valid content security measures”. It also warns that listing paths in robots.txt exposes them publicly, so a Disallow line can point curious visitors straight at the thing you wanted hidden.

Anthropic adds a related warning: blocking its IP addresses may not guarantee an opt-out, because doing so stops the bot reading your robots.txt in the first place. Keep robots.txt readable and add other measures alongside it.

Which fetchers may ignore it

OpenAI says robots.txt rules may not apply to ChatGPT-User because users initiate the action. Perplexity says Perplexity-User generally ignores robots.txt rules, and Amazon says Amzn-User may not follow all directives.

There’s also the question of bots that don’t identify themselves honestly. Cloudflare reported in August 2025 that it saw Perplexity changing its user agent and network when blocked, and removed Perplexity as a verified bot; that is Cloudflare’s account of its own tests, not a finding we could confirm independently.

How to enforce a block

For real enforcement, block at the WAF, CDN or server. Helen Pollitt, Head of SEO at Getty Images, wrote in Search Engine Journal that the higher up the stack you go the better: WAF if you can, CDN if you can’t, server as a last resort.

Cloudflare’s AI Crawl Control lets you set allow or block rules for individual crawlers and track which ones violate your robots.txt. Perplexity’s documentation recommends combining user-agent matching with its published IP ranges if you want to allow its bots through a firewall.

What should you check before using a CDN preset?

Presets can catch more than you intend. Cloudflare said in a 1 July 2026 post that, from 15 September 2026, multi-purpose crawlers that combine search with training would be treated according to all their behaviours, so Googlebot, Applebot and Bingbot would be blocked for customers who had chosen to block training.

We haven’t tested how that plays out on live sites. If you use Cloudflare, open the AI settings and confirm that the search crawlers you rely on are still allowed before you switch on any training block.

Related: What Is llms.txt and Does Your Site Need One?

How Do You Check Which AI Bots Visit Your Site?

Search your server or CDN logs for the bot names, then check any IP address against the vendor’s published list, because a user-agent string can be faked. Ten minutes in the logs tells you more than any list of bots.

What to look for in the logs

Count hits per bot name, then look at the response codes. A bot you’ve blocked in robots.txt that still returns 200 on content pages is worth investigating.

This command counts hits per AI bot in a standard access log:

grep -oE "GPTBot|OAI-SearchBot|ChatGPT-User|ClaudeBot|Claude-SearchBot|Claude-User|PerplexityBot|Perplexity-User|Amzn-SearchBot|Amzn-User|Amazonbot|CCBot" access.log | sort | uniq -c | sort -rn

Applebot-Extended and Google-Extended won’t appear, because neither has a user agent of its own. Apple says Applebot-Extended doesn’t crawl, and Google says Google-Extended crawls with existing user agents.

How to verify a bot is genuine

OpenAI, Perplexity, Common Crawl and Apple publish IP lists or reverse DNS checks, and Anthropic’s help article links to its own IP list. Use them, because spoofing a user-agent string costs nothing.

Bot How to verify
OAI-SearchBot, GPTBot, ChatGPT-User IP lists at openai.com/searchbot.json, gptbot.json and chatgpt-user.json
Claude bots IP list at claude.com/crawling/bots.json
PerplexityBot, Perplexity-User IP lists at perplexity.com/perplexitybot.json and perplexity-user.json
CCBot Reverse DNS ending crawl.commoncrawl.org, or the list at index.commoncrawl.org/ccbot.json
Applebot Reverse DNS in the *.applebot.apple.com domain, or Apple’s published CIDR file

A worked example

Imagine a log line like this one on a small business site:

203.0.113.7 - - [03/Oct/2026:09:14:22 +0000] "GET /guides/ HTTP/1.1" 200 18452 "-" "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.4; +https://openai.com/gptbot"

The user agent says GPTBot, but if 203.0.113.7 isn’t in OpenAI’s gptbot.json list, you’re looking at something pretending to be GPTBot. The 200 response would also tell you any robots.txt rule is being ignored, which is the cue to add a firewall rule rather than another robots.txt line.

Related: How Do You Measure AI Search Visibility?

Should You Block AI Crawlers or Allow Them?

For most businesses that want leads, allow the search and retrieval bots and make a deliberate choice about the training bots. Publishers whose content is the product have a stronger case for blocking more.

Type of site Training bots Search and retrieval bots Reason
Local service business or agency Your call Allow Being named in an answer is the whole point
Lead generation or comparison site Your call Allow Visitors arrive from answers, so blocking trades visits for little
Paywalled research or licensed data Block Allow free pages only The content is the asset
Original news or reporting Block, and consider a firewall rule Allow selectively Licensing power matters more than reach

Allow if you rely on being found

Picture a Cheltenham accountancy firm whose best leads come from people asking an assistant for a local specialist. Blocking OAI-SearchBot or PerplexityBot there trades visibility for very little.

The same logic runs through generative engine optimisation, which assumes the bots can get in.

Block if your content is the asset

If you sell research behind a paywall, opting out of training is a reasonable position. Do it knowing it’s a request, and back it with firewall rules if the content really is sensitive.

Review on a schedule

Bot names, versions and policies change, and so do Google’s controls, as the Search Console toggle shows. Diarise a review each quarter and compare your robots.txt against the vendor pages.

Crawler access is the first thing to rule out in any AI search or SEO project, because nothing else works if the bots can’t get in.

If you’d like an outside view of what AI engines can reach and say about your brand, get an AI visibility audit from £500.

Training crawlersSearch and retrieval bots
ExamplesGPTBot, ClaudeBot, AmazonbotOAI-SearchBot, Claude-SearchBot, PerplexityBot
What they feedModel trainingSearch indexes and cited answers
If you block themContent opted out of trainingYou can drop out of AI answers
Typical decisionA content and licensing choiceA visibility choice
Two kinds of AI crawler, two different decisions

FAQs

Will blocking AI crawlers hurt my Google rankings?

Not if you only block the AI vendors' bots. Google says Google-Extended doesn't impact a site's inclusion in Google Search nor is it used as a ranking signal, and OpenAI, Anthropic and Perplexity run separate bots from Googlebot. Blocking Googlebot itself would remove you from Search, so never do that by accident.

How long does a robots.txt change take to apply?

OpenAI and Amazon both say about 24 hours, and Perplexity says up to 24 hours. Anthropic doesn't publish a figure. The robots.txt standard itself says crawlers shouldn't use a cached copy for more than 24 hours unless the file is unreachable.

Do I need to block each subdomain separately?

Yes. Anthropic says to apply the rule for every subdomain you want to opt out, and Amazon says its bots read robots.txt at the host level and honour the rules under each host. Treat blog.example.com and www.example.com as two separate files.

What happens if my robots.txt returns an error?

It depends on the error. Under RFC 9309, a 4xx response means the crawler may access anything, while a 5xx response or an unreachable server means it must assume a complete disallow. A broken server can therefore hide your whole site from well-behaved bots.

Can I slow AI crawlers down instead of blocking them?

Only with some of them. Anthropic says its bots support the non-standard Crawl-delay directive, while Apple and Amazon both say their crawlers don't follow it. Neither OpenAI nor Perplexity mention it in their crawler documentation.

Should I block CCBot?

CCBot belongs to Common Crawl, a non-profit that maintains an open repository of web crawl data that anyone can analyse. If you don't want to be in that repository, Common Crawl's own documentation gives a Disallow rule for CCBot, and we found no statement from Common Crawl about what happens to pages already archived.

What is Cloudflare's pay per crawl?

It's a Cloudflare feature that lets AI crawlers pay to access your content. Cloudflare's documentation lists it as a private beta, so most sites can't use it yet.

Sources

  1. Overview of OpenAI Crawlers, OpenAI
  2. Does Anthropic crawl data from the web, and how can site owners block the crawler?, Anthropic Help Center
  3. Perplexity Crawlers, Perplexity
  4. Google's common crawlers, Google Search Central
  5. Search generative AI control, Search Console Help
  6. New opportunities, control and insights for website owners, Google, The Keyword
  7. AI features and your website, Google Search Central
  8. CMA secures fairer deal for publishers and improves Google search services in UK, GOV.UK, Competition and Markets Authority
  9. Google search publisher conduct requirement, GOV.UK, Competition and Markets Authority
  10. About Applebot, Apple Support
  11. About Amazonbot, Amazon
  12. CCBot, Common Crawl
  13. RFC 9309: Robots Exclusion Protocol, IETF
  14. Your site, your rules: new AI traffic options for all customers, Cloudflare
  15. Perplexity is using stealth, undeclared crawlers to evade website no-crawl directives, Cloudflare
  16. AI Crawl Control, Cloudflare Docs
  17. Should I Block AI Crawlers At Robots.txt Or Server Level? (Ask An SEO), Search Engine Journal, Helen Pollitt

Related services

Want this done for you? Let’s talk.

Tell us about your site and your goals. You get a clear scope and price before any work starts.

Talk to a specialistGet an AI visibility audit