Guide · Search & AI Visibility

robots.txt for AI: GPTBot, ClaudeBot, PerplexityBot and what to block

Setting up robots.txt for AI is no longer a matter of saying “yes” or “no” to artificial intelligence: each company runs two or three different bots, and some train models while others cite you in their answers. Here you will find a verified table of user agents, a robots.txt builder and a clear rule of thumb for blocking whatever you like without vanishing from ChatGPT, Perplexity or Google.

By Updated: 10 min read

The basics

What is robots.txt and why does it matter now with AI?

Short answer

A robots.txt for AI is your robots.txt file configured to decide which artificial intelligence bots can crawl your site, and for what. There are three kinds of AI bot: training crawlers (GPTBot, ClaudeBot), search crawlers (OAI-SearchBot, PerplexityBot) and agents acting on a user’s request (ChatGPT-User). You can block training without losing visibility; block the search crawlers and you disappear from their answers.

robots.txt is a plain-text file that sits at the root of your domain (yourdomain.co.uk/robots.txt). It works in groups: a User-agent line says which bot you are addressing, and beneath it go the Disallow rules (keep out of here) and Allow rules (this is fine). Since September 2022 it has been an official standard, the IETF’s RFC 9309.

Until recently, thinking about Googlebot and Bingbot was enough. Today OpenAI, Anthropic, Perplexity, Meta, Apple, Amazon and Mistral all publish their own AI crawlers, each with a different purpose. The question is no longer “do I let AI in, yes or no?” but which use of your content you accept: models being trained on it, answers citing it, or assistants reading it when a user asks. That, at heart, is what configuring a robots.txt for AI means: deciding bot by bot, not blindly.

So before you block AI bots, it pays to know what each one does. Blocking the wrong one protects nothing and can cost you visibility.

That decision has a direct impact on your AI SEO: a misplaced block can wipe you from ChatGPT Search or Perplexity without anyone telling you. If you want the full picture of how visibility is earned in search engines and answer engines, I cover it in the guide on what is GEO and how it relates to SEO and AEO.

robots.txt is not a lock

RFC 9309 says so itself: its rules “are not a form of access authorization”. It is a request that reputable bots honour, not a technical barrier. And because the file is public, listing “secret” paths in it makes them easier to find.

The key concept

GPTBot, OAI-SearchBot and ChatGPT-User: why are there three bots per company?

Because each one does a different job, and blocking it has different consequences. This is the central idea behind any robots.txt for AI. OpenAI is the clearest example: there is not one ChatGPT user agent but three, and its official documentation puts it in writing: “each setting is independent of the others”.

1. Training crawlers

GPTBot, ClaudeBot, Meta-ExternalAgent, CCBot and MistralAI-Training collect text that may end up in the training data of future models. Blocking them does not remove you from any answer today: it simply signals that you do not want your future content feeding those models. What has already been collected is not deleted.

2. Search crawlers

OAI-SearchBot, Claude-SearchBot, PerplexityBot and Meta-WebIndexer build the index that the citations and links in search-backed answers come from. OpenAI is blunt about it: sites that opt out of OAI-SearchBot “will not be shown in ChatGPT search answers”. These are the ones you almost never want to block.

3. Agents acting on a user’s request

ChatGPT-User, Claude-User, Perplexity-User and Meta-ExternalFetcher do not crawl on their own initiative: they open a specific URL when someone asks the assistant to read it. Here is the catch: OpenAI warns that, because these actions are initiated by a user, “robots.txt rules may not apply”, and Perplexity, Meta and Amazon say something similar. Anthropic, by contrast, states that all its bots respect robots.txt.

The special cases: Google, Apple and Microsoft

Google has no separate AI bot. AI Overviews and AI Mode are part of Search and draw on what Googlebot crawls. Google-Extended is not a crawler but a control token: it decides whether content already crawled can be used to train Gemini and for grounding, and Google says it does not affect your inclusion or ranking in Search. Apple follows the same model with Applebot and Applebot-Extended. Microsoft has no separate AI bot either: Copilot relies on the Bingbot index, and its webmaster guidelines, rewritten in February 2026 control use in Copilot through the noarchive meta tag.

Key points

  • Blocking a training bot (GPTBot, ClaudeBot) does not cost you visibility in AI answers.
  • Blocking a search bot (OAI-SearchBot, PerplexityBot, Claude-SearchBot) does remove you from its answers.
  • You cannot opt out of AI Overviews with robots.txt without leaving Google: both depend on Googlebot.
  • Google-Extended and Applebot-Extended are control tokens: they only affect training.
  • Several user agents warn that they may ignore robots.txt: for a real block you need your CDN or WAF.

Verified October 2026

Which AI crawlers read robots.txt? A verified user-agent table

This is the table I use to configure a robots.txt for AI. Every row has been checked against each company’s official documentation as of 9 October 2026; where the provider gives no detail or there is doubt, I say so. Names change often, so check the date if you land here a few months from now.

AI user agents and what happens if you block them (October 2026)
CriterionCompanyWhat it doesRespects robots.txt?If you block it…
GPTBotOpenAITrainingYesYour content is not used to train OpenAI models. No effect on ChatGPT Search.
OAI-SearchBotOpenAISearch (ChatGPT Search)YesChanges take about 24 hoursYou do not appear in ChatGPT’s search answers.
ChatGPT-UserOpenAIUser actions (and custom GPTs)Not alwaysRules “may not apply”, says OpenAIIt may still visit when a user asks it to.
ClaudeBotAnthropicTrainingYesAlso supports Crawl-delayYour future content stays out of Claude’s training.
Claude-SearchBotAnthropicSearchYesClaude does not index your site: less visibility in its search-backed answers.
Claude-UserAnthropicUser actionsYesAnthropic says all its bots respect itClaude cannot read your page when a user asks it to.
PerplexityBotPerplexitySearchDoes not train models, says PerplexityYes, says PerplexityChallenged by Cloudflare in August 2025You are not cited or linked in Perplexity.
Perplexity-UserPerplexityUser actionsGenerally notAccording to its own documentationIt may keep visiting at users’ request.
GooglebotGoogleSearch: Google, AI Overviews and AI ModeYesYou disappear from Google, including AI Overviews and AI Mode.
Google-ExtendedGoogleControl token (does not crawl)Gemini training and groundingYes (it only exists in robots.txt)Gemini does not train on your site. You stay in Search and AI Overviews.
ApplebotAppleSearch: Siri, Spotlight and SafariYesYou drop out of Siri, Spotlight and Safari search results.
Applebot-ExtendedAppleControl token (does not crawl)Apple model trainingYes (it only exists in robots.txt)Apple does not train on your site. You still appear in its search results.
BingbotMicrosoftSearch: Bing and CopilotYesYou drop out of Bing and Copilot. To limit Copilot only, use the noarchive meta tag.
CCBotCommon CrawlOpen archiveWidely used to train third-party modelsYes (official opt-out method)Your pages are left out of new Common Crawl crawls.
Meta-ExternalAgentMetaTraining (and product improvement)YesYour site is excluded from Meta model training.
Meta-WebIndexerMetaMeta AI searchYesMeta AI does not cite or link you in its answers.
Meta-ExternalFetcherMetaUser actionsNot always“May bypass” robots.txt, says MetaIt may still visit at a user’s request.
BytespiderByteDanceData collectionNo public official documentationUnconfirmedBlock it in your CDN or WAF as well if you want to be sure.
AmazonbotAmazonProduct improvementMay be used to train Amazon modelsYesAmazon does not use your site for those purposes. Does not support Crawl-delay.
Amzn-SearchBotAmazonSearch (Alexa and others)YesYou do not appear in Amazon’s search experiences.
Amzn-UserAmazonUser actionsNot alwaysMay not follow every directive, says AmazonIt may still visit to answer a live request.
MistralAI-TrainingMistralTrainingYesMistral does not use your site to train its models.
MistralAI-IndexMistralSearchVibe assistant (formerly Le Chat)Not specifiedIts search engine does not index your site for answers in Vibe.
MistralAI-UserMistralUser actionsYes, says Mistralrobots.txt decides which sites it may visitVibe cannot open your page at a user’s request.

Recent changes worth keeping on your radar:

  • OpenAI also documents OAI-AdsBot, which only visits pages submitted as ads in ChatGPT to check their safety. It is not a general crawler.
  • Meta now separates training (Meta-ExternalAgent), Meta AI search (Meta-WebIndexer) and user requests (Meta-ExternalFetcher).
  • Mistral publishes three agents: MistralAI-Training, MistralAI-Index and MistralAI-User. Its assistant is now called Vibe (formerly Le Chat).
  • Google lists Google-Agent among its user-triggered fetchers; it is used by agents hosted on Google’s infrastructure to browse the web. Like the rest of that family, it “generally ignores” robots.txt rules.
  • Perplexity: on 4 August 2025 Cloudflare accused it of using undeclared crawlers to evade blocks and removed it from its list of verified bots. Perplexity rejected the accusation. That is why I flag it with a caveat in the table.
  • ByteDance publishes no official documentation for Bytespider that I have been able to verify. In 2024 Cloudflare identified it as the AI bot making the most requests on its network.

24 h

How long ChatGPT Search can take to reflect a change to your robots.txt

OpenAI, crawler documentation

4

Companies in this table that warn their user agent may not respect robots.txt

OpenAI, Perplexity, Meta and Amazon docs

500 KiB

Minimum robots.txt size a crawler must be able to parse

IETF, RFC 9309 (2022)

Tool

How do you block AI bots with a robots.txt disallow rule?

With one group per bot (or several bots in the same group) and a Disallow: / rule. The standard lets you stack several User-agent lines that share the same rules, which keeps the file shorter:

robots.txt
User-agent: GPTBot
User-agent: ClaudeBot
Disallow: /

What many people look up as robots.txt disallow is exactly this rule: Disallow: / closes the whole site to the bots in that group. It can also be partial: Disallow: /clients/ blocks just that folder. If several rules match a URL, the most specific one wins, meaning the longest path.

The most common mistake: groups do not add up

A bot that finds a group with its own name ignores the User-agent: * group entirely. If you write a User-agent: GPTBot group with Allow: /, GPTBot will no longer see your general Disallow rules for /wp-admin/ or /cart/. If you give a bot its own group, repeat in it any rules you need.

To avoid mistakes, use this robots.txt for AI builder. Pick a quick preset or switch bots on and off one by one: the file is generated instantly, new lines are highlighted and the panel tells you what happens to your visibility.

robots.txt builder

Generated in your browser
Training

Training9/9 allowed

They collect text to train models. Blocking them does not cost you visibility in answers.

User-triggered

User-triggered6/6 allowed

They open a page when someone asks the assistant to. The flagged ones warn that they may ignore robots.txt.

robots.txt
# Generated at josegalan.dev/en/guides
# Check your private paths before publishing it
 
User-agent: *
Allow: /
 
Sitemap: https://www.yourdomain.co.uk/sitemap.xml

What this set-up means

  • You will appear in ChatGPT SearchYes
  • You will appear in AI Overviews and AI ModeYes
  • You will appear in CopilotYes
  • You will appear in PerplexityYes
  • Claude can index you for its searchesYes
  • Your content can be used for training9 of 9 training crawlers allowedYes

You will appear in ChatGPT Search: yes. You will appear in AI Overviews and AI Mode: yes. You will appear in Copilot: yes. You will appear in Perplexity: yes. Claude can index you for its searches: yes. Your content can be used for training: yes.

The builder deliberately leaves Googlebot, Bingbot and Applebot out of the “Block all AI” preset: blocking them also takes you out of Google, Bing and Apple’s search results. If that is genuinely what you want, you can switch them off by hand.

Watch out

Is robots.txt enough? Cloudflare, CDNs and WAFs

No. You can have a perfect robots.txt for AI and still be blocking AI without realising it, or the other way round, because your CDN or firewall gets the final say before your file does.

The most important case is Cloudflare, which serves a very large share of the web. On 1 July 2025 it announced that it blocks AI crawlers accessing content without permission by default: since then, every new domain is asked whether it wants to allow them. The one-click option to block them had existed since 2024, and today it is managed bot by bot in AI Crawl Control.

Cloudflare can also serve a managed robots.txt that adds the line Content-Signal: search=yes, ai-train=no. This is its Content Signals Policy (24 September 2025): a way of expressing usage preferences that, as Cloudflare itself acknowledges, some companies may ignore. It is not currently part of the RFC 9309 standard.

Check your CDN before blaming your content

If your site never appears in ChatGPT Search or Perplexity, the first thing to check is that OAI-SearchBot and PerplexityBot are not getting a 403 error from your CDN, your WAF or a security plugin. I see it more often than you might think.

And if you need a bot genuinely kept out, the tool is your WAF or CDN, not robots.txt: it is the only layer that can stop an agent that ignores the file or identifies itself under another name.

Step by step

How can you check which AI bots visit your site?

In five steps, with no paid tools. It will take you less than an hour the first time.

  1. Open your live robots.txt

    Visit yourdomain.co.uk/robots.txt in your browser: that is what bots see, not what you think you configured in your CMS. In Google Search Console, the robots.txt report tells you which version Google has read and whether it contains errors. If your CDN serves a managed robots.txt, it may not match your site’s own file.
  2. Look for the user agents in your server logs

    Download the access log (Apache or Nginx, from your hosting panel) and count each bot’s visits with the command below. You will see who comes in, how often and to which URLs.
  3. Check they are who they claim to be

    User agents can be spoofed. Check the IP against the official lists: OpenAI publishes gptbot.json, searchbot.json and chatgpt-user.json; Anthropic, claude.com/crawling/bots.json; Common Crawl, ccbot.json; and Amazon and Mistral publish their own lists.
  4. Review your CDN and WAF

    In Cloudflare, look at AI Crawl Control and the security events: which bots are being blocked and with which status code. If OAI-SearchBot or PerplexityBot are getting a 403, your robots.txt makes no difference.
  5. Check the outcome in the answers

    Ask ChatGPT with search, Perplexity and Copilot about your service and your brand, and note whether they cite you and with which URL. Repeat every month with the same questions to track progress.
terminal
# Which AI bots have visited your site? (Apache or Nginx log)
BOTS="GPTBot|OAI-SearchBot|ChatGPT-User|Claude|Perplexity|CCBot|Bytespider"
BOTS="$BOTS|Meta-External|Meta-WebIndexer|Amazonbot|Amzn-|Applebot|MistralAI"
grep -oiE "[A-Za-z-]*($BOTS)[A-Za-z-]*" access.log | sort | uniq -c | sort -rn

Common questions

Frequently asked questions about robots.txt for AI

What is the ChatGPT user agent in robots.txt?

ChatGPT has three user agents, not one: GPTBot collects content to train OpenAI’s models, OAI-SearchBot decides whether you appear in ChatGPT’s search answers, and ChatGPT-User visits a page when a user asks it to read one. That is why blocking the ChatGPT user agent for training (GPTBot) does not remove you from ChatGPT: OpenAI states that each setting is independent of the others.

What is ClaudeBot and should I block it?

ClaudeBot is the crawler Anthropic uses to collect web content that may be used to train its Claude models. Blocking it signals that your future content should be kept out of training, but it does not remove you from Claude’s search-backed answers, which rely on Claude-SearchBot and Claude-User. If your site is mostly commercial content, letting it through usually pays off.

How do I block AI bots in robots.txt without leaving Google?

To block AI bots without leaving Google, list each AI company’s training, search and user crawlers in robots.txt, add Google-Extended and Applebot-Extended, and leave Googlebot, Bingbot and Applebot allowed. The catch: AI Overviews and AI Mode use Googlebot, so robots.txt cannot take you out of them without taking you out of Google; you can only limit snippets with nosnippet or max-snippet.

Does Google-Extended remove me from AI Overviews?

No. Google-Extended is not a crawler but a control token: it tells Google whether content it already crawls can be used to train future Gemini models and for grounding in other Google systems. Google says it does not affect inclusion or ranking in Search, and AI Overviews and AI Mode are part of Search, which is governed by Googlebot.

Do AI crawlers always respect robots.txt?

There is no guarantee. The RFC 9309 standard itself says robots.txt rules are not a form of access authorisation. The major providers state that their automated AI crawlers respect it, but OpenAI, Perplexity, Meta and Amazon warn that their user-triggered agents may not follow it. If you need a real block, set it up in your CDN or WAF as well.

How long does a change to robots.txt for AI take to apply?

Usually about a day. RFC 9309 asks crawlers not to use a cached copy older than 24 hours; OpenAI mentions around 24 hours for ChatGPT Search, and Perplexity, Meta and Amazon give similar timeframes. Bear in mind that a change to your robots.txt for AI does not erase what has already been crawled or what a model has already learnt; it only affects what happens from then on.

Does a robots.txt disallow for AI bots hurt my Google SEO?

Not if you do it properly. Blocking AI bots such as GPTBot, ClaudeBot, PerplexityBot or Google-Extended does not affect your Google rankings; Google explicitly says Google-Extended is not a ranking signal. The risk lies in mistakes, such as a robots.txt disallow that is too broad (Disallow: / under User-agent: *) or blocking Googlebot by accident. Check the file in Search Console’s robots.txt report after every change.

Verification

Sources

All consulted on 9 October 2026. Bot names and behaviour change often: if something does not add up, the official source takes precedence.

This article was created with the help of AI and reviewed by José Galán. I take great care over every post and every translation, but the odd mistake can still slip through. If you find one, write to me: you will be helping me improve.

Search & AI Visibility

Want to know whether AI can read (and cite) your site?

I review your robots.txt, your CDN and your logs, check which AI bots actually visit you and measure whether you appear in ChatGPT, Perplexity, Copilot and AI Overviews. With a clear report on what to change and why.

See the service

Written by

Agent-as-a-Service Designer at AllHub and Search & AI Visibility consultant. I design internal AI agents and help companies show up in Google and AI engines, with reproducible measurement.

LinkedInBackground

Shall we talk?

Get recommended by AI and/or put an agent to work for you.

Tell me what you want to achieve: fill in the form or email me at info@josegalan.dev.

Basic data protection information. Controller: José Galán. Purpose: to reply to your enquiry. Legal basis: your consent, which you can withdraw at any time. Recipients: no data is shared with third parties unless required by law; Google (hosting and form delivery) acts as data processor. Your rights: access, rectification, erasure, objection, restriction and portability. More information in the privacy policy.

      robots.txt for AI: which AI bots to block | José Galán