AI Crawlers Explained: GPTBot, OAI-SearchBot, PerplexityBot
AI companies run separate crawlers for AI search, for user requests and for model training, and you can treat each differently in robots.txt. Allow the search crawlers if you want assistants to cite you; blocking the training crawlers is a legitimate choice.
AI crawlers are web robots operated by AI companies; each declares its own name (user agent) so site owners can allow or block it individually in robots.txt.
- Search crawlers (OAI-SearchBot, PerplexityBot, Claude-SearchBot) decide whether assistants can cite you.
- Training crawlers (GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, CCBot) decide whether your content may be used to train models.
- User-triggered fetchers (ChatGPT-User, Perplexity-User, Claude-User) visit a page because a person asked, and may not follow robots.txt.
- Blocking everything in one rule also blocks the crawlers that bring you recommendations.
Which AI crawlers exist, and what is each one for?
| Crawler | Operator | Purpose (as documented) | robots.txt |
|---|---|---|---|
| OAI-SearchBot | OpenAI | Surfaces websites in ChatGPT's search features | Controls it; changes can take about 24 hours |
| GPTBot | OpenAI | Crawls content that may be used for training | Controls it |
| ChatGPT-User | OpenAI | Visits a page when a user asks ChatGPT or a custom GPT | May not apply, because a user initiates it |
| PerplexityBot | Perplexity | Surfaces and links sites in Perplexity search; not used for foundation model training | Respects it |
| Perplexity-User | Perplexity | Supports user requests inside Perplexity | Generally ignores it, because a user asked |
| ClaudeBot | Anthropic | Collects content that could contribute to model training | Respects it |
| Claude-SearchBot | Anthropic | Improves search result quality | Respects it |
| Claude-User | Anthropic | Supports user-initiated requests | Anthropic says its bots honour robots.txt directives |
| Google-Extended | Token controlling use for Gemini training and grounding; no separate user agent | A robots.txt token; no effect on Search | |
| Applebot-Extended | Apple | Controls whether Applebot-crawled data trains Apple models; it does not crawl | A robots.txt token |
| CCBot | Common Crawl | Builds the Common Crawl dataset | Respects it |
Sources: OpenAI, Perplexity, Anthropic, Google, Apple and Common Crawl. Names and behaviour change, so check the vendor page before relying on a detail.
Should you allow or block AI crawlers?
It depends on what you want:
- You want assistants to recommend you: allow the search crawlers (OAI-SearchBot, PerplexityBot, Claude-SearchBot) and make sure Googlebot can index your pages, since Google says AI Overviews and AI Mode need a page that is indexed and eligible to show with a snippet (Google).
- You do not want your content used for training: disallow GPTBot, ClaudeBot, Google-Extended, Applebot-Extended and CCBot. OpenAI says disallowing GPTBot "indicates a site's content should not be used in training", and Apple says pages that disallow Applebot-Extended can still appear in its search results.
- Both: allow the search crawlers and block the training ones. The two groups are independent.
What does a sensible robots.txt look like?
This example lets the search crawlers in, blocks the training crawlers and keeps private paths private. Adapt the paths to your site:
User-agent: *
Disallow: /admin/
User-agent: OAI-SearchBot
User-agent: PerplexityBot
User-agent: Claude-SearchBot
Allow: /
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: CCBot
Disallow: /Note that a crawler with its own group follows only that group, so copy any private-path rules you still want (such as Disallow: /admin/) into groups that allow access.
Why does a robots.txt rule not always stop a visit?
User-triggered fetchers act for a person. OpenAI says robots.txt rules "may not apply" to ChatGPT-User, and Perplexity says Perplexity-User generally ignores them. If you must keep such visits out entirely, use a firewall or authentication, not robots.txt. Anthropic also notes that blocking by IP address may not reliably prevent crawling, and publishes an IP list to verify its bots.
How do you check what your site allows today?
- Run the AI crawler checker with your domain. It shows each crawler's verdict and the exact robots.txt line that decides it.
- If a search crawler is blocked, copy the fixed robots.txt it generates, which keeps your private paths disallowed.
- Upload the file to the root of your site and re-check after about a day; OpenAI says OAI-SearchBot can take about 24 hours to adjust.
- Add an llms.txt so crawlers that can read your site find your key pages quickly.
- Run the AI-readiness audit for the rest of the technical basics, then the free AI Visibility Check to see whether assistants name you for your market.
Frequently asked questions
What is the difference between GPTBot and OAI-SearchBot?
OAI-SearchBot surfaces websites in ChatGPT's search features, so blocking it keeps you out of those results. GPTBot crawls content that may be used to train OpenAI's models; blocking it does not remove you from ChatGPT search.
Will blocking Google-Extended hurt my Google rankings?
No. Google states that Google-Extended does not impact a site's inclusion in Google Search and is not a ranking signal. It only controls use of crawled content for training future Gemini models and for grounding.
Should a small business block AI training crawlers?
That is your call. Blocking training crawlers keeps your content out of future model training, and it does not stop the search crawlers that let assistants cite you, as long as you allow those separately.
Why does my firewall block AI crawlers even though robots.txt allows them?
robots.txt only states your policy. Firewalls, CDNs and bot-protection tools can block crawlers regardless of it. If the checker shows access allowed but assistants still cannot read your site, check your bot-protection settings.
Does ChatGPT recommend your business?
Run the free AI Visibility Check: 10 real buyer questions, ChatGPT and Perplexity, your score against 3 competitors. No signup.
Keep reading
- How to Get Your Business Recommended by ChatGPT: A ChecklistA practical checklist for small businesses to become visible in ChatGPT, Perplexity and Gemini answers: access, clarity, proof and weekly measurement.
- How AI Assistants Choose Which Businesses to RecommendHow ChatGPT, Perplexity, Gemini and Claude find businesses to recommend: what the engines document, what they keep secret, and what you can influence.
- llms.txt: What It Is and How to Write One for Your Sitellms.txt is a proposed plain-markdown file that tells AI assistants what your site is about. What the spec says, a worked example and a six-step guide.
- What Is GEO? Generative Engine Optimization ExplainedGEO (Generative Engine Optimization) is how you get ChatGPT, Perplexity and Gemini to name your business. A plain-English guide and a six-step start.