The basics
What is robots.txt and why does it matter now with AI?
Short answer
A robots.txt for AI is your robots.txt file configured to decide which artificial intelligence bots can crawl your site, and for what. There are three kinds of AI bot: training crawlers (GPTBot, ClaudeBot), search crawlers (OAI-SearchBot, PerplexityBot) and agents acting on a user’s request (ChatGPT-User). You can block training without losing visibility; block the search crawlers and you disappear from their answers.
robots.txt is a plain-text file that sits at the root of your domain (yourdomain.co.uk/robots.txt). It works in groups: a User-agent line says which bot you are addressing, and beneath it go the Disallow rules (keep out of here) and Allow rules (this is fine). Since September 2022 it has been an official standard, the IETF’s RFC 9309.
Until recently, thinking about Googlebot and Bingbot was enough. Today OpenAI, Anthropic, Perplexity, Meta, Apple, Amazon and Mistral all publish their own AI crawlers, each with a different purpose. The question is no longer “do I let AI in, yes or no?” but which use of your content you accept: models being trained on it, answers citing it, or assistants reading it when a user asks. That, at heart, is what configuring a robots.txt for AI means: deciding bot by bot, not blindly.
So before you block AI bots, it pays to know what each one does. Blocking the wrong one protects nothing and can cost you visibility.
That decision has a direct impact on your AI SEO: a misplaced block can wipe you from ChatGPT Search or Perplexity without anyone telling you. If you want the full picture of how visibility is earned in search engines and answer engines, I cover it in the guide on what is GEO and how it relates to SEO and AEO.
robots.txt is not a lock
The key concept
GPTBot, OAI-SearchBot and ChatGPT-User: why are there three bots per company?
Because each one does a different job, and blocking it has different consequences. This is the central idea behind any robots.txt for AI. OpenAI is the clearest example: there is not one ChatGPT user agent but three, and its official documentation puts it in writing: “each setting is independent of the others”.
1. Training crawlers
GPTBot, ClaudeBot, Meta-ExternalAgent, CCBot and MistralAI-Training collect text that may end up in the training data of future models. Blocking them does not remove you from any answer today: it simply signals that you do not want your future content feeding those models. What has already been collected is not deleted.
2. Search crawlers
OAI-SearchBot, Claude-SearchBot, PerplexityBot and Meta-WebIndexer build the index that the citations and links in search-backed answers come from. OpenAI is blunt about it: sites that opt out of OAI-SearchBot “will not be shown in ChatGPT search answers”. These are the ones you almost never want to block.
3. Agents acting on a user’s request
ChatGPT-User, Claude-User, Perplexity-User and Meta-ExternalFetcher do not crawl on their own initiative: they open a specific URL when someone asks the assistant to read it. Here is the catch: OpenAI warns that, because these actions are initiated by a user, “robots.txt rules may not apply”, and Perplexity, Meta and Amazon say something similar. Anthropic, by contrast, states that all its bots respect robots.txt.
The special cases: Google, Apple and Microsoft
Google has no separate AI bot. AI Overviews and AI Mode are part of Search and draw on what Googlebot crawls. Google-Extended is not a crawler but a control token: it decides whether content already crawled can be used to train Gemini and for grounding, and Google says it does not affect your inclusion or ranking in Search. Apple follows the same model with Applebot and Applebot-Extended. Microsoft has no separate AI bot either: Copilot relies on the Bingbot index, and its webmaster guidelines, rewritten in February 2026 control use in Copilot through the noarchive meta tag.
Key points
- Blocking a training bot (GPTBot, ClaudeBot) does not cost you visibility in AI answers.
- Blocking a search bot (OAI-SearchBot, PerplexityBot, Claude-SearchBot) does remove you from its answers.
- You cannot opt out of AI Overviews with robots.txt without leaving Google: both depend on Googlebot.
- Google-Extended and Applebot-Extended are control tokens: they only affect training.
- Several user agents warn that they may ignore robots.txt: for a real block you need your CDN or WAF.
Verified October 2026
Which AI crawlers read robots.txt? A verified user-agent table
This is the table I use to configure a robots.txt for AI. Every row has been checked against each company’s official documentation as of 9 October 2026; where the provider gives no detail or there is doubt, I say so. Names change often, so check the date if you land here a few months from now.
| Criterion | Company | What it does | Respects robots.txt? | If you block it… |
|---|---|---|---|---|
GPTBot | OpenAI | Training | Yes | Your content is not used to train OpenAI models. No effect on ChatGPT Search. |
OAI-SearchBot | OpenAI | Search (ChatGPT Search) | YesChanges take about 24 hours | You do not appear in ChatGPT’s search answers. |
ChatGPT-User | OpenAI | User actions (and custom GPTs) | Not alwaysRules “may not apply”, says OpenAI | It may still visit when a user asks it to. |
ClaudeBot | Anthropic | Training | YesAlso supports Crawl-delay | Your future content stays out of Claude’s training. |
Claude-SearchBot | Anthropic | Search | Yes | Claude does not index your site: less visibility in its search-backed answers. |
Claude-User | Anthropic | User actions | YesAnthropic says all its bots respect it | Claude cannot read your page when a user asks it to. |
PerplexityBot | Perplexity | SearchDoes not train models, says Perplexity | Yes, says PerplexityChallenged by Cloudflare in August 2025 | You are not cited or linked in Perplexity. |
Perplexity-User | Perplexity | User actions | Generally notAccording to its own documentation | It may keep visiting at users’ request. |
Googlebot | Search: Google, AI Overviews and AI Mode | Yes | You disappear from Google, including AI Overviews and AI Mode. | |
Google-Extended | Control token (does not crawl)Gemini training and grounding | Yes (it only exists in robots.txt) | Gemini does not train on your site. You stay in Search and AI Overviews. | |
Applebot | Apple | Search: Siri, Spotlight and Safari | Yes | You drop out of Siri, Spotlight and Safari search results. |
Applebot-Extended | Apple | Control token (does not crawl)Apple model training | Yes (it only exists in robots.txt) | Apple does not train on your site. You still appear in its search results. |
Bingbot | Microsoft | Search: Bing and Copilot | Yes | You drop out of Bing and Copilot. To limit Copilot only, use the noarchive meta tag. |
CCBot | Common Crawl | Open archiveWidely used to train third-party models | Yes (official opt-out method) | Your pages are left out of new Common Crawl crawls. |
Meta-ExternalAgent | Meta | Training (and product improvement) | Yes | Your site is excluded from Meta model training. |
Meta-WebIndexer | Meta | Meta AI search | Yes | Meta AI does not cite or link you in its answers. |
Meta-ExternalFetcher | Meta | User actions | Not always“May bypass” robots.txt, says Meta | It may still visit at a user’s request. |
Bytespider | ByteDance | Data collectionNo public official documentation | Unconfirmed | Block it in your CDN or WAF as well if you want to be sure. |
Amazonbot | Amazon | Product improvementMay be used to train Amazon models | Yes | Amazon does not use your site for those purposes. Does not support Crawl-delay. |
Amzn-SearchBot | Amazon | Search (Alexa and others) | Yes | You do not appear in Amazon’s search experiences. |
Amzn-User | Amazon | User actions | Not alwaysMay not follow every directive, says Amazon | It may still visit to answer a live request. |
MistralAI-Training | Mistral | Training | Yes | Mistral does not use your site to train its models. |
MistralAI-Index | Mistral | SearchVibe assistant (formerly Le Chat) | Not specified | Its search engine does not index your site for answers in Vibe. |
MistralAI-User | Mistral | User actions | Yes, says Mistralrobots.txt decides which sites it may visit | Vibe cannot open your page at a user’s request. |
Recent changes worth keeping on your radar:
- OpenAI also documents
OAI-AdsBot, which only visits pages submitted as ads in ChatGPT to check their safety. It is not a general crawler. - Meta now separates training (
Meta-ExternalAgent), Meta AI search (Meta-WebIndexer) and user requests (Meta-ExternalFetcher). - Mistral publishes three agents:
MistralAI-Training,MistralAI-IndexandMistralAI-User. Its assistant is now called Vibe (formerly Le Chat). - Google lists
Google-Agentamong its user-triggered fetchers; it is used by agents hosted on Google’s infrastructure to browse the web. Like the rest of that family, it “generally ignores” robots.txt rules. - Perplexity: on 4 August 2025 Cloudflare accused it of using undeclared crawlers to evade blocks and removed it from its list of verified bots. Perplexity rejected the accusation. That is why I flag it with a caveat in the table.
- ByteDance publishes no official documentation for
Bytespiderthat I have been able to verify. In 2024 Cloudflare identified it as the AI bot making the most requests on its network.
24 h
How long ChatGPT Search can take to reflect a change to your robots.txt
OpenAI, crawler documentation
4
Companies in this table that warn their user agent may not respect robots.txt
OpenAI, Perplexity, Meta and Amazon docs
500 KiB
Minimum robots.txt size a crawler must be able to parse
IETF, RFC 9309 (2022)
Tool
How do you block AI bots with a robots.txt disallow rule?
With one group per bot (or several bots in the same group) and a Disallow: / rule. The standard lets you stack several User-agent lines that share the same rules, which keeps the file shorter:
User-agent: GPTBot
User-agent: ClaudeBot
Disallow: /What many people look up as robots.txt disallow is exactly this rule: Disallow: / closes the whole site to the bots in that group. It can also be partial: Disallow: /clients/ blocks just that folder. If several rules match a URL, the most specific one wins, meaning the longest path.
The most common mistake: groups do not add up
User-agent: * group entirely. If you write a User-agent: GPTBot group with Allow: /, GPTBot will no longer see your general Disallow rules for /wp-admin/ or /cart/. If you give a bot its own group, repeat in it any rules you need.To avoid mistakes, use this robots.txt for AI builder. Pick a quick preset or switch bots on and off one by one: the file is generated instantly, new lines are highlighted and the panel tells you what happens to your visibility.
The builder deliberately leaves Googlebot, Bingbot and Applebot out of the “Block all AI” preset: blocking them also takes you out of Google, Bing and Apple’s search results. If that is genuinely what you want, you can switch them off by hand.
My recommendation
Which robots.txt for AI do I recommend for a small business?
For most small businesses, the best starting point is to let every search and user bot through and to take your time over training. If your site is mostly commercial content (services, case studies, pricing, guides), models knowing about your brand tends to work in your favour.
# Recommended robots.txt for most small businesses (October 2026)
# Goal: maximum visibility in search engines and in AI.
User-agent: *
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php
Disallow: /cart/
Disallow: /my-account/
Disallow: /*?s=
# Crawler with no public official documentation
User-agent: Bytespider
Disallow: /
# Optional: if you do NOT want models trained on your content,
# remove the hashes from this block. It does not affect your
# visibility in ChatGPT Search, Perplexity or Google.
# User-agent: GPTBot
# User-agent: ClaudeBot
# User-agent: Google-Extended
# User-agent: Applebot-Extended
# User-agent: CCBot
# User-agent: Meta-ExternalAgent
# User-agent: MistralAI-Training
# Disallow: /
Sitemap: https://www.yourdomain.co.uk/sitemap.xmlAdapt the private paths to your CMS: the ones in the example are typical of WordPress with WooCommerce. The training block is commented out so that you only switch it on if it suits you. I have left out Amazonbot because it mixes product improvement and training; add it if you want to be strict.
When is blocking training worth it?
robots.txt decides what bots can read. If you also want to suggest which pages models should read first, there is a complementary proposal: I explain what it is, who actually uses it and how much to expect from it in the guide llms.txt: what it is and how to create one.
Watch out
Is robots.txt enough? Cloudflare, CDNs and WAFs
No. You can have a perfect robots.txt for AI and still be blocking AI without realising it, or the other way round, because your CDN or firewall gets the final say before your file does.
The most important case is Cloudflare, which serves a very large share of the web. On 1 July 2025 it announced that it blocks AI crawlers accessing content without permission by default: since then, every new domain is asked whether it wants to allow them. The one-click option to block them had existed since 2024, and today it is managed bot by bot in AI Crawl Control.
Cloudflare can also serve a managed robots.txt that adds the line Content-Signal: search=yes, ai-train=no. This is its Content Signals Policy (24 September 2025): a way of expressing usage preferences that, as Cloudflare itself acknowledges, some companies may ignore. It is not currently part of the RFC 9309 standard.
Check your CDN before blaming your content
And if you need a bot genuinely kept out, the tool is your WAF or CDN, not robots.txt: it is the only layer that can stop an agent that ignores the file or identifies itself under another name.
Step by step
How can you check which AI bots visit your site?
In five steps, with no paid tools. It will take you less than an hour the first time.
Open your live robots.txt
Visityourdomain.co.uk/robots.txtin your browser: that is what bots see, not what you think you configured in your CMS. In Google Search Console, the robots.txt report tells you which version Google has read and whether it contains errors. If your CDN serves a managed robots.txt, it may not match your site’s own file.Look for the user agents in your server logs
Download the access log (Apache or Nginx, from your hosting panel) and count each bot’s visits with the command below. You will see who comes in, how often and to which URLs.Check they are who they claim to be
User agents can be spoofed. Check the IP against the official lists: OpenAI publishesgptbot.json,searchbot.jsonandchatgpt-user.json; Anthropic,claude.com/crawling/bots.json; Common Crawl,ccbot.json; and Amazon and Mistral publish their own lists.Review your CDN and WAF
In Cloudflare, look at AI Crawl Control and the security events: which bots are being blocked and with which status code. If OAI-SearchBot or PerplexityBot are getting a 403, your robots.txt makes no difference.Check the outcome in the answers
Ask ChatGPT with search, Perplexity and Copilot about your service and your brand, and note whether they cite you and with which URL. Repeat every month with the same questions to track progress.
# Which AI bots have visited your site? (Apache or Nginx log)
BOTS="GPTBot|OAI-SearchBot|ChatGPT-User|Claude|Perplexity|CCBot|Bytespider"
BOTS="$BOTS|Meta-External|Meta-WebIndexer|Amazonbot|Amzn-|Applebot|MistralAI"
grep -oiE "[A-Za-z-]*($BOTS)[A-Za-z-]*" access.log | sort | uniq -c | sort -rnCommon questions
Frequently asked questions about robots.txt for AI
What is the ChatGPT user agent in robots.txt?
ChatGPT has three user agents, not one: GPTBot collects content to train OpenAI’s models, OAI-SearchBot decides whether you appear in ChatGPT’s search answers, and ChatGPT-User visits a page when a user asks it to read one. That is why blocking the ChatGPT user agent for training (GPTBot) does not remove you from ChatGPT: OpenAI states that each setting is independent of the others.
What is ClaudeBot and should I block it?
ClaudeBot is the crawler Anthropic uses to collect web content that may be used to train its Claude models. Blocking it signals that your future content should be kept out of training, but it does not remove you from Claude’s search-backed answers, which rely on Claude-SearchBot and Claude-User. If your site is mostly commercial content, letting it through usually pays off.
How do I block AI bots in robots.txt without leaving Google?
To block AI bots without leaving Google, list each AI company’s training, search and user crawlers in robots.txt, add Google-Extended and Applebot-Extended, and leave Googlebot, Bingbot and Applebot allowed. The catch: AI Overviews and AI Mode use Googlebot, so robots.txt cannot take you out of them without taking you out of Google; you can only limit snippets with nosnippet or max-snippet.
Does Google-Extended remove me from AI Overviews?
No. Google-Extended is not a crawler but a control token: it tells Google whether content it already crawls can be used to train future Gemini models and for grounding in other Google systems. Google says it does not affect inclusion or ranking in Search, and AI Overviews and AI Mode are part of Search, which is governed by Googlebot.
Do AI crawlers always respect robots.txt?
There is no guarantee. The RFC 9309 standard itself says robots.txt rules are not a form of access authorisation. The major providers state that their automated AI crawlers respect it, but OpenAI, Perplexity, Meta and Amazon warn that their user-triggered agents may not follow it. If you need a real block, set it up in your CDN or WAF as well.
How long does a change to robots.txt for AI take to apply?
Usually about a day. RFC 9309 asks crawlers not to use a cached copy older than 24 hours; OpenAI mentions around 24 hours for ChatGPT Search, and Perplexity, Meta and Amazon give similar timeframes. Bear in mind that a change to your robots.txt for AI does not erase what has already been crawled or what a model has already learnt; it only affects what happens from then on.
Does a robots.txt disallow for AI bots hurt my Google SEO?
Not if you do it properly. Blocking AI bots such as GPTBot, ClaudeBot, PerplexityBot or Google-Extended does not affect your Google rankings; Google explicitly says Google-Extended is not a ranking signal. The risk lies in mistakes, such as a robots.txt disallow that is too broad (Disallow: / under User-agent: *) or blocking Googlebot by accident. Check the file in Search Console’s robots.txt report after every change.
Verification
Sources
All consulted on 9 October 2026. Bot names and behaviour change often: if something does not add up, the official source takes precedence.
- OpenAI: Overview of OpenAI Crawlers (GPTBot, OAI-SearchBot, ChatGPT-User, OAI-AdsBot).
- Anthropic: Does Anthropic crawl data from the web, and how can site owners block the crawler?
- Perplexity: Perplexity Crawlers
- Google Search Central: Google’s common crawlers (Googlebot, Google-Extended).
- Google Search Central: AI features and your website
- Google Search Central: Google’s user-triggered fetchers (Google-Agent).
- Apple: About Applebot
- Bing Webmaster Blog: Announcing new options for webmasters to control usage of their content in Bing Chat (22 September 2023)
- Bing Webmaster Guidelines and Search Engine Journal’s summary of the February 2026 changes
- Common Crawl: CCBot
- Meta for Developers: Meta Web Crawlers
- Amazon: Amazonbot, Amzn-SearchBot and Amzn-User
- Mistral AI: Robots
- IETF: RFC 9309, Robots Exclusion Protocol (September 2022)
- Cloudflare: press release of 1 July 2025 and AI Crawl Control documentation
- Cloudflare: Content Signals Policy (24 September 2025)
- Cloudflare: report on Perplexity’s undeclared crawlers (4 August 2025)
- Tom’s Guide: Perplexity’s response to Cloudflare (August 2025)
- heise: Mistral’s chatbot is now called Vibe
- Cloudflare: Declaring your AIndependence (3 July 2024)
This article was created with the help of AI and reviewed by José Galán. I take great care over every post and every translation, but the odd mistake can still slip through. If you find one, write to me: you will be helping me improve.
Search & AI Visibility
Want to know whether AI can read (and cite) your site?
I review your robots.txt, your CDN and your logs, check which AI bots actually visit you and measure whether you appear in ChatGPT, Perplexity, Copilot and AI Overviews. With a clear report on what to change and why.
See the service