KNOWLEDGE BASE · GUIDE

AI Crawlers Explained: GPTBot, ClaudeBot, PerplexityBot, Google-Extended

GPTBot, ClaudeBot, PerplexityBot and Google-Extended are the crawlers that decide whether AI answer engines can read your site — allow them on public pages and block them on private ones, or you will not exist in that engine's answers. This page gives the copy-paste rules and what each bot actually does.

Who they are

Crawler Operator Reads for Allowed means
GPTBot OpenAI Model training material Content may inform model knowledge
OAI-SearchBot / ChatGPT-User OpenAI Search index / user-triggered fetch Answers can quote and link your pages
ClaudeBot / Claude-SearchBot / Claude-User Anthropic Training / search / user-fetch Same split as OpenAI's trio
PerplexityBot / Perplexity-User Perplexity Ongoing search index / on-demand Citations in Perplexity answers
Google-Extended Google Gemini/Vertex grounding, not web search Blocked = excluded from grounding, Search unaffected
Applebot-Extended Apple Apple Intelligence Excluded = no Siri/Apple answers

Roles per the audit's crawler-access checklist (skillpilotadvisory.ai/geo and /pillar/become-the-answer, live 11 Oct 2026).

The copy-paste robots.txt block (adopted in our own /robots.txt)

User-agent: GPTBot
Allow: /
Disallow: /my-products

User-agent: OAI-SearchBot
Allow: /

User-agent: ChatGPT-User
Allow: /

User-agent: ClaudeBot
Allow: /
Disallow: /my-products

User-agent: Claude-SearchBot
Allow: /

User-agent: Claude-User
Allow: /

User-agent: PerplexityBot
Allow: /

User-agent: Perplexity-User
Allow: /

User-agent: Google-Extended
Allow: /

User-agent: Applebot-Extended
Allow: /

Sitemap: https://skillpilotadvisory.ai/sitemap.xml

Plus the site-wide block: User-agent: * with Allow: / and Disallow: for /account, /audit-account, /business-case-account, /my-products, /api/. Engines that ignore named rules fall back to the * block. (Live in our /robots.txt.)

The two rules that matter

  1. Block the workspace, not the marketing site. Private accounts stay disallowed; public facts stay crawlable.
  2. Allowing is not endorsing. Allowing a training crawler is a business decision — decide it consciously, write it down, and date the decision. Allowing search/user bots (OAI-SearchBot, ChatGPT-User, Claude-User, Perplexity-User) is the low-controversy, high-return half: that is what gets you quoted.

Sources and context

  • "If a bot is blocked, you do not exist in that engine's answers" (/pillar/become-the-answer, live 11 Oct 2026).
  • Most small and mid-size sites run crawler rules "written for search, not for answers" — no llms.txt, thin entity markup, bots blocked by default (skillpilotadvisory.ai/geo, fetched 10 Oct 2026).
  • Crawler access is point 2 of the 7-point GEO method and checks 1-6 of the 40-check audit (same pillar).

FAQ

Will allowing GPTBot let AI train on my content? GPTBot's role is model training; if that is unwanted, block it but allow OAI-SearchBot and ChatGPT-User so search answers can still cite you.
Does blocking Google-Extended hurt my Google rankings? No — it grounds Gemini/Vertex answers, not web search indexing.
Do I need per-bot blocks for private pages? The greedy * block covers private paths; add the per-bot Disallow only where the crawler honours it strictly.
How do I test? Ask each engine about your site, and check server logs for the bot's user agent.

START, THINK, WORK, GROW WITH AI

Your next move.
Not another maybe.