Alle artikelen
5 oktober 2026

How AI crawlers work: GPTBot, ClaudeBot, PerplexityBot and your robots.txt choices

GPTBot, ClaudeBot and PerplexityBot are not one bot each. Here is what every AI crawler does and how to control them in robots.txt.

Fiber patch cables plugged into rack servers with green port lights, the hardware AI crawlers request pages from

OpenAI, Anthropic and Perplexity each run more than one crawler. Some train a model. Some power a live search feature. Some only fetch a page because a person asked an assistant to look at it. You control each one separately in robots.txt, by its own user agent name. As of September 2022, robots.txt is a formal internet standard, RFC 9309, but it remains a voluntary convention: a crawler stays out only if its operator chooses to honor the file.

What AI crawlers are and why they are not search engine crawlers

Traditional search crawlers like Googlebot do one job: find pages and index them for search results. AI crawlers split that job into three distinct categories, and each category gets its own bot name.

Node map diagram with three parts: ChatGPT-User, Claude-User and Perplexity-User
Node map: What AI crawlers are and why they are not search engine crawlers.

Training crawlers harvest content to improve a foundation model. GPTBot and ClaudeBot fall here, alongside Google-Extended and Bytespider. Search and index crawlers power an AI product's own answer or search surface, separate from training: OAI-SearchBot, Claude-SearchBot and PerplexityBot all do this job. User-triggered fetchers retrieve a single page because a live user asked an assistant to look at it right now, which industry guidance treats as distinct from automatic crawling: ChatGPT-User, Claude-User and Perplexity-User.

Why does the distinction matter? Because the three categories are controlled independently. You can disallow a company's training bot while still allowing its search bot to fetch your pages. That keeps your content out of model training while it stays eligible for citation in an AI answer. We cover the citation side in more detail in how to get cited by ChatGPT, Perplexity and Google AI Overviews.

How it works in practice

Each company publishes its own user agent names, its own IP ranges, and its own robots.txt syntax. None of the three treat "block the AI company" as a single switch.

OpenAI's four bots

OpenAI documents four separate crawlers. GPTBot is used to make OpenAI's generative AI foundation models more useful and safe, which is training. OAI-SearchBot surfaces websites in ChatGPT's search results, a separate function from training. OAI-AdsBot validates the safety of pages submitted as ads, and OpenAI states explicitly that data it collects is not used to train generative AI foundation models. ChatGPT-User handles actions a live user triggers inside ChatGPT, such as pasting a URL, and is not used for automatic crawling. Each bot publishes its own IP range as a JSON file. OpenAI notes a robots.txt change can take about 24 hours to reach its systems.

To block GPTBot only, while leaving search access open:

User-agent: GPTBot
Disallow: /

User-agent: OAI-SearchBot
Allow: /

Anthropic's three bots

Anthropic runs ClaudeBot, Claude-User and Claude-SearchBot, and each does a different job. ClaudeBot collects web content for training. Claude-User retrieves a page only when a live user asks Claude a question about it. Claude-SearchBot indexes content for Claude's search features. To rate-limit rather than fully block a bot, Anthropic's own guidance gives this syntax:

User-agent: ClaudeBot
Crawl-delay: 1

To block it outright, replace the second line with Disallow: /. Anthropic publishes its crawler IP ranges at claude.com/crawling/bots.json, though it also notes that blocking by IP address alone isn't reliable since ranges change.

Perplexity's two bots

Perplexity runs the same split in miniature: PerplexityBot is its search and index crawler, and Perplexity-User is the on-demand fetcher triggered by a live query, the same role ChatGPT-User and Claude-User play for their companies. As with the others, each name gets its own User-agent line.

Getting these settings right is part of the same technical SEO and AEO work we do for clients. See our SEO and AEO work for how the audit fits together.

Tradeoffs and edge cases

robots.txt only works if the crawler operator reads it and chooses to comply. That's worth saying plainly, because the gap between publishing a robots.txt policy and actually following it became a public dispute in 2025.

robots.txt is a convention, not a lock

The file has no enforcement mechanism. It originated in 1994 and only became a formal standard, RFC 9309, in September 2022, nearly three decades later. A compliant crawler checks the file before fetching. A noncompliant one doesn't have to.

The Perplexity dispute

Cloudflare reported on August 4, 2025 that after a site blocked Perplexity's declared crawler through robots.txt or network rules, Perplexity's traffic kept arriving anyway, through undeclared crawlers that rotated IP addresses and user agent strings. Cloudflare said it observed the pattern across tens of thousands of domains, amounting to millions of requests per day, and removed Perplexity from its verified bot program as a result. Perplexity disputed the characterization, calling the report a publicity stunt and arguing Cloudflare conflated legitimate user-triggered fetches with autonomous crawling. Either way, the episode shows something simple: a robots.txt rule is a request a compliant operator honors, not a wall every operator respects.

Who actually blocks what

Ahrefs analyzed roughly 140 million sites, publishing its results on May 21, 2025. GPTBot was blocked by 5.89% of sites, ClaudeBot by 5.74%, Google-Extended by 5.71%, and PerplexityBot by 5.61%. ClaudeBot's block rate grew 32.67% over the prior year, the fastest growth of any crawler Ahrefs tracked. robots.txt is one lever among several. Pairing it with an AI-specific file is covered in llms.txt explained.

Frequently asked questions

Does blocking GPTBot in robots.txt remove my site from ChatGPT answers?

No. GPTBot and OAI-SearchBot are separate OpenAI crawlers with separate robots.txt tokens. Disallowing GPTBot keeps your content out of model training, while OAI-SearchBot can still fetch pages for ChatGPT's search feature. Training access and citation access are controlled independently.

Is robots.txt legally binding on AI companies?

No. robots.txt, standardized as RFC 9309 in September 2022, is a voluntary convention. Compliant crawlers like GPTBot and ClaudeBot check it before fetching, but it carries no enforcement mechanism. That's part of why Cloudflare's August 2025 report on Perplexity's undeclared crawlers became a story at all.

What is the difference between ClaudeBot, Claude-User and Claude-SearchBot?

ClaudeBot collects web content for training Anthropic's models. Claude-User fetches a specific page only when a live user asks Claude about it. Claude-SearchBot indexes content for Claude's search features. Each has its own User-agent token, so a site can block training while allowing the other two.

How do I block every known AI crawler at once?

There is no single wildcard token that reliably covers every AI bot, since new ones appear regularly. Add explicit User-agent blocks for the current major names (GPTBot, ClaudeBot, PerplexityBot, Google-Extended, Bytespider, CCBot and their user-triggered counterparts) and revisit the list periodically: the number of distinct AI bots roughly doubled between August 2023 and December 2024.

Why did Cloudflare remove Perplexity from its verified bot program?

Cloudflare reported on August 4, 2025 that after being blocked by a site's robots.txt or network rules, Perplexity's traffic continued under undeclared, stealth crawlers that rotated IP addresses and user agents, across tens of thousands of domains. Perplexity disputed the characterization, saying Cloudflare conflated legitimate user-triggered fetches with autonomous crawling.

Blocking training while keeping citation access, bot by bot, is the practical version of this decision. The harder limit is that robots.txt is a request, not a wall, and the list of bots asking keeps changing. Pairing that choice with a broader answer engine optimization strategy makes sure blocking training never quietly costs you citations too.