KNOWLEDGE BASE · GUIDE
AI Crawlers Explained: GPTBot, ClaudeBot, PerplexityBot, Google-Extended
GPTBot, ClaudeBot, PerplexityBot and Google-Extended are the crawlers that decide whether AI answer engines can read your site — allow them on public pages and block them on private ones, or you will not exist in that engine's answers. This page gives the copy-paste rules and what each bot actually does.
Who they are
| Crawler | Operator | Reads for | Allowed means |
|---|---|---|---|
| GPTBot | OpenAI | Model training material | Content may inform model knowledge |
| OAI-SearchBot / ChatGPT-User | OpenAI | Search index / user-triggered fetch | Answers can quote and link your pages |
| ClaudeBot / Claude-SearchBot / Claude-User | Anthropic | Training / search / user-fetch | Same split as OpenAI's trio |
| PerplexityBot / Perplexity-User | Perplexity | Ongoing search index / on-demand | Citations in Perplexity answers |
| Google-Extended | Gemini/Vertex grounding, not web search | Blocked = excluded from grounding, Search unaffected | |
| Applebot-Extended | Apple | Apple Intelligence | Excluded = no Siri/Apple answers |
Roles per the audit's crawler-access checklist (skillpilotadvisory.ai/geo and /pillar/become-the-answer, live 11 Oct 2026).
The copy-paste robots.txt block (adopted in our own /robots.txt)
User-agent: GPTBot
Allow: /
Disallow: /my-products
User-agent: OAI-SearchBot
Allow: /
User-agent: ChatGPT-User
Allow: /
User-agent: ClaudeBot
Allow: /
Disallow: /my-products
User-agent: Claude-SearchBot
Allow: /
User-agent: Claude-User
Allow: /
User-agent: PerplexityBot
Allow: /
User-agent: Perplexity-User
Allow: /
User-agent: Google-Extended
Allow: /
User-agent: Applebot-Extended
Allow: /
Sitemap: https://skillpilotadvisory.ai/sitemap.xml
Plus the site-wide block: User-agent: * with Allow: / and Disallow: for /account, /audit-account, /business-case-account, /my-products, /api/. Engines that ignore named rules fall back to the * block. (Live in our /robots.txt.)
The two rules that matter
- Block the workspace, not the marketing site. Private accounts stay disallowed; public facts stay crawlable.
- Allowing is not endorsing. Allowing a training crawler is a business decision — decide it consciously, write it down, and date the decision. Allowing search/user bots (OAI-SearchBot, ChatGPT-User, Claude-User, Perplexity-User) is the low-controversy, high-return half: that is what gets you quoted.
Sources and context
- "If a bot is blocked, you do not exist in that engine's answers" (/pillar/become-the-answer, live 11 Oct 2026).
- Most small and mid-size sites run crawler rules "written for search, not for answers" — no llms.txt, thin entity markup, bots blocked by default (skillpilotadvisory.ai/geo, fetched 10 Oct 2026).
- Crawler access is point 2 of the 7-point GEO method and checks 1-6 of the 40-check audit (same pillar).
FAQ
Will allowing GPTBot let AI train on my content? GPTBot's role is model training; if that is unwanted, block it but allow OAI-SearchBot and ChatGPT-User so search answers can still cite you.
Does blocking Google-Extended hurt my Google rankings? No — it grounds Gemini/Vertex answers, not web search indexing.
Do I need per-bot blocks for private pages? The greedy * block covers private paths; add the per-bot Disallow only where the crawler honours it strictly.
How do I test? Ask each engine about your site, and check server logs for the bot's user agent.