Key Takeaways
- Know which crawler you are blocking. Each AI company runs separate crawlers for training, for search, and for fetching a page a person asked about. A block on one does not touch the others.
- Keep the search crawlers open if you want to be recommended. OAI-SearchBot, Claude-SearchBot and PerplexityBot are how those assistants find pages to cite. Blocking them takes you out of the answers you are trying to get into.
- Treat training crawlers as a separate decision. GPTBot, ClaudeBot and Google-Extended govern training use. Blocking them is a reasonable choice and, by the owners' own documentation, does not remove you from search.
- Remember robots.txt is a request, not a lock. Reputable crawlers follow it by choice. Anything genuinely private belongs behind a login, not behind a line in a text file.
- Read your own file before you assume anything. Add /robots.txt to your site address and read it. Some hosting and security settings add AI blocks you never chose, and that takes two minutes to catch.
Somebody told you to block the AI bots. Somebody else told you AI is how clients find coaches now. Both pieces of advice arrive with confidence and they point in opposite directions.
The good news is that they are mostly about different crawlers. Each AI company runs more than one, for more than one job, and the file that controls them, robots.txt, lets you answer each job separately.
Here is whether AI crawlers obey robots.txt, the three jobs they do, every named crawler side by side, whether you should block any of them, and how to check what your own site is saying right now.
Do AI crawlers obey robots.txt?
Mostly yes, by their owners' own account. The crawlers that roam the web on a schedule are documented as following robots.txt. The fetchers that act when a person asks about a specific page are the exception, and two of the three big ones say so plainly.
Robots.txt is a plain text file at the root of your site that tells crawlers which parts they may read. It works on trust. Google's introduction to robots.txt says the instructions cannot enforce crawler behavior and that it is up to the crawler to obey them. It also warns that the file is not a way to keep a page out of Google.
So the honest answer is layered. The named, reputable AI crawlers say they comply, and the file is still a request rather than a wall.
Three jobs, three different crawlers
Every major assistant separates training, search, and on-request fetching into different crawlers with different names. The name you block decides which of those three things you switch off.
OpenAI's crawler documentation lists GPTBot for training, OAI-SearchBot for surfacing sites in ChatGPT search, and ChatGPT-User for actions a person starts. Of that last one, it says that because the actions are initiated by a user, robots.txt rules may not apply. Anthropic's help page on how site owners can block its crawlers describes the same split, ClaudeBot, Claude-SearchBot and Claude-User, says all three respect robots.txt, and warns that disabling Claude-SearchBot may reduce your site's visibility in its search results.
- Owner says it follows robots.txt
- Owner says robots.txt may not apply
Training
Your call
- GPTBot OpenAI
- ClaudeBot Anthropic
- Google-Extended Google
Blocking removes future training reads. It does not remove you from search answers.
Search index
Keep open
- OAI-SearchBot OpenAI
- Claude-SearchBot Anthropic
- PerplexityBot Perplexity
Blocking removes you from the pages that assistant can cite when it searches.
Fetch on request
Keep open
- ChatGPT-User OpenAI
- Claude-User Anthropic
- Perplexity-User Perplexity
Runs when a person asks about a page. Two of three owners say robots.txt may not apply.
Perplexity's guide to its crawlers completes the picture. PerplexityBot surfaces and links sites in its search results and is not used for model training. Perplexity-User visits pages when a person asks a question, and Perplexity says that because a user requested the fetch, it generally ignores robots.txt.
Every named AI crawler, side by side
Nine names cover the three assistants most coaches care about plus Google's training control. The last column is the part that matters for a business that wants to be recommended.
| Crawler | Owner | Job | Follows robots.txt? | If you block it |
|---|---|---|---|---|
| GPTBot | OpenAI | Training | Yes | Not used to train future models |
| OAI-SearchBot | OpenAI | ChatGPT search | Yes | Not surfaced in ChatGPT search results |
| ChatGPT-User | OpenAI | Fetch on a person's request | May not apply | Block may be ignored |
| ClaudeBot | Anthropic | Training | Yes | Not used to train future models |
| Claude-SearchBot | Anthropic | Claude search | Yes | May reduce visibility in Claude search results |
| Claude-User | Anthropic | Fetch on a person's request | Yes | Claude cannot open the page when asked |
| PerplexityBot | Perplexity | Perplexity search | Yes | Not surfaced in Perplexity results |
| Perplexity-User | Perplexity | Fetch on a person's request | Generally ignores it | Block likely ignored |
| Google-Extended | Gemini training control | Yes | No effect on Google Search inclusion or ranking |
One absence is deliberate. Google AI Overviews are not on this list because they draw on the ordinary Google Search index, which is read by Googlebot. Blocking Googlebot to stay out of AI Overviews would also take you out of Google Search, which is almost never the trade a coach wants.
Should you block AI crawlers?
Block training crawlers if you want to; keep search and on-request crawlers open if you want to be recommended. Those are two separate decisions, and most bad advice comes from treating them as one.
The case for keeping search crawlers open is simple. When a buyer asks ChatGPT, Claude or Perplexity to find a coach, those assistants often search the live web and cite what they find. A page their search crawler cannot read is a page they cannot quote. Blocking OAI-SearchBot to protect your content, then wondering why ChatGPT never names you, is a common and entirely self-inflicted problem.
The case for blocking training crawlers is also real. Your writing is yours, and not every business wants it inside a model. The owners' documentation says blocking GPTBot, ClaudeBot or Google-Extended does not remove you from their search results, so the choice is genuinely separate.
The damaging admission: there is a cost we cannot measure precisely. What a model already knows about you, without searching, came from training data. A business that blocks every training crawler may be slightly less familiar to assistants when they answer from memory. Nobody has published a clean number for that effect, and we will not invent one. If your pages are mostly public-facing marketing, leaving training open costs you little. If they are paid material, blocking it is sensible.
How to check what your site is telling them
Type your site address followed by /robots.txt into a browser and read what comes back. It takes two minutes and it is the only way to know for sure.
Search the page for GPTBot, OAI-SearchBot, ClaudeBot, Claude-SearchBot and PerplexityBot. A line reading Disallow: / under any of those names blocks it from the whole site. Also look for a group headed User-agent: * with Disallow: /, which blocks every crawler that follows the file. If none of the names appear, the general rules apply to them.
Look out for rules you did not write. Some website builders, hosting plans and security services offer a switch that blocks AI crawlers, and a well-meaning web person may have flipped it. If you find a block you did not choose, change the setting where it was made, then allow about a day for ChatGPT search to catch up, which is the figure OpenAI gives.
Access is the floor, not the ceiling. Once crawlers can read your site, what they find still has to be current and specific; see how to keep a page current so AI quotes it right. Two other files come up in the same conversation: what llms.txt is and whether it helps, and what IndexNow does for getting changes noticed faster.
If you want to see whether the doors being open is translating into being named, start with the free check. It asks ChatGPT 25 real buyer questions from your niche at $0 and emails a plain-English readout: your score, who got named in your place, and the one fix to start with. It runs as soon as you ask for it, and the report lands within the hour between 7am and 5pm Central, or by 10am the next morning outside those hours.
Questions we hear the most
Do AI crawlers obey robots.txt?
The ones that crawl on their own schedule do, according to their owners: GPTBot, OAI-SearchBot, ClaudeBot, Claude-SearchBot and PerplexityBot are all documented as following it. Fetchers that act on a person's request are different, and two of them say the file may not apply.
Should I block AI crawlers from my website?
Not the search crawlers, if you want to be recommended. Blocking OAI-SearchBot, Claude-SearchBot or PerplexityBot removes your pages from what those assistants can cite. Blocking training crawlers is a separate, reasonable choice that does not remove you from search.
Does blocking GPTBot stop ChatGPT from recommending me?
Not by itself, according to OpenAI. GPTBot is the training crawler, and ChatGPT search uses OAI-SearchBot, which OpenAI documents separately. A site can block GPTBot and still allow OAI-SearchBot.
Does blocking Google-Extended remove me from Google?
No. Google says Google-Extended does not affect a site's inclusion in Google Search and is not a ranking signal. It controls whether content Google crawls may be used to train future Gemini models.
Is robots.txt enough to keep a page private?
No. Google's own guidance says robots.txt cannot enforce crawler behavior and is not a mechanism for keeping a page out of Google. Anything private should sit behind a login or carry a noindex rule instead.
How long does a robots.txt change take to apply?
For ChatGPT search, OpenAI says it can take about 24 hours for its systems to adjust after you update the file. Other assistants do not publish a figure, so check again after a few days rather than assuming.
Can the free check tell me if my robots.txt is blocking AI?
No, and it is worth being plain about that. The free check asks ChatGPT 25 real buyer questions from your niche at $0 and reports whether you are named. Reading your robots.txt is a two-minute job you can do yourself.
Does the $500 audit look at crawler access?
It shows the effect rather than the file: 125 checks across ChatGPT, Claude, Gemini, Perplexity and Google AI Overviews, so an assistant that never names you stands out against the others. It lands by this time tomorrow.
Your buyers are already asking. Find out what AI tells them.
The free check asks ChatGPT 25 real buyer questions about your niche and sends you the report within the hour, at $0. If it shows you are already getting named everywhere, we will say so in plain words.
Get my free check →