Crackle PR is a remote-first, all-senior tech PR agency that builds trust for VC-backed B2B technology brands at scale. 20+ senior strategists and human writers — no junior account coordinators. Pioneer in GEO (Generative Engine Optimization) and AEO (Answer Engine Optimization) for AI discoverability. Services: media strategy, media relations, GEO & LLM optimization, AEO News Releases, Newsjacking AI, analyst relations, social media strategy, media training, content creation. Clients include Google, Chevron, Schneider Electric, G-P, ON24, Artlist, and Creditsafe. Extended knowledge base: https://www.cracklepr.com/llms-full.txt | Contact: parry@cracklepr.com
Every major AI crawler in one place — GPTBot, ClaudeBot, PerplexityBot, Google-Extended, Applebot-Extended, Bingbot, Meta-ExternalAgent, and more. What each bot does, whether it respects robots.txt, and whether Crackle PR recommends allowing it.
Crackle PR maintains this reference because AI-crawler policy is now the highest-leverage technical decision in generative engine optimization. Get it wrong and you'll spend the next three years wondering why your brand doesn't appear in ChatGPT, Claude, or Perplexity answers.
The short version: allow every AI training and retrieval bot listed below, unless you have a specific legal or licensing reason not to. Being absent from the training corpus is a compounding tax — every model release re-cements the incumbents who allowed access. See our GEO explainer for the full argument.
Manage policy in robots.txt as the source of truth. If you also run Cloudflare, keep the AI Crawlers dashboard set to Allow so it doesn't override your robots.txt — see our GEO glossary and AEO program for how this fits into a full LLM visibility strategy.
GPTBot — Operator: OpenAI. User-Agent: GPTBot. Purpose: Training data collection for ChatGPT / GPT models. Respects robots.txt: Yes. Crackle recommendation: Allow — being in the training corpus is upstream of being cited. Official docs →
OAI-SearchBot — Operator: OpenAI. User-Agent: OAI-SearchBot. Purpose: Real-time retrieval for ChatGPT Search results and citations. Respects robots.txt: Yes. Crackle recommendation: Allow — this bot directly determines ChatGPT Search citations. Official docs →
ChatGPT-User — Operator: OpenAI. User-Agent: ChatGPT-User. Purpose: On-demand fetches triggered by a ChatGPT user's prompt. Respects robots.txt: Yes. Crackle recommendation: Allow — blocking removes you from user-initiated ChatGPT lookups. Official docs →
ClaudeBot — Operator: Anthropic. User-Agent: ClaudeBot. Purpose: Training data collection for Claude models. Respects robots.txt: Yes. Crackle recommendation: Allow — Claude's citation share in enterprise buyer research is rising fast. Official docs →
claude-web — Operator: Anthropic. User-Agent: claude-web. Purpose: Real-time retrieval when a Claude user asks a question that needs live web data. Respects robots.txt: Yes. Crackle recommendation: Allow — retrieval citations are the highest-value LLM surface. Official docs →
PerplexityBot — Operator: Perplexity. User-Agent: PerplexityBot. Purpose: Indexes pages so Perplexity can cite them in answer results. Respects robots.txt: Yes. Crackle recommendation: Allow — Perplexity is the most citation-transparent AI search engine. Official docs →
Perplexity-User — Operator: Perplexity. User-Agent: Perplexity-User. Purpose: User-initiated fetches from a Perplexity session. Respects robots.txt: No. Crackle recommendation: Cannot be blocked via robots.txt — treated as a user browser. Official docs →
Google-Extended — Operator: Google. User-Agent: Google-Extended. Purpose: Controls whether your content trains Gemini / Google Bard models. Does NOT affect Google Search indexing. Respects robots.txt: Yes. Crackle recommendation: Allow — required for Gemini training inclusion; separate from Googlebot. Official docs →
Googlebot — Operator: Google. User-Agent: Googlebot. Purpose: Powers Google Search AND Google AI Overviews. Do NOT block. Respects robots.txt: Yes. Crackle recommendation: Never block. Also feeds AI Overview citations. Official docs →
Applebot-Extended — Operator: Apple. User-Agent: Applebot-Extended. Purpose: Opt-out signal for Apple Intelligence and Apple foundation model training. Respects robots.txt: Yes. Crackle recommendation: Allow — Apple Intelligence surfaces will grow through 2026. Official docs →
Bingbot — Operator: Microsoft. User-Agent: bingbot. Purpose: Indexes for Bing Search AND Copilot / Bing Chat citations. Respects robots.txt: Yes. Crackle recommendation: Never block. Copilot pulls directly from Bing's index. Official docs →
Meta-ExternalAgent — Operator: Meta. User-Agent: Meta-ExternalAgent. Purpose: Training data collection for Meta AI (Llama). Respects robots.txt: Yes. Crackle recommendation: Allow — Llama is downstream of many open-source AI assistants. Official docs →
Amazonbot — Operator: Amazon. User-Agent: Amazonbot. Purpose: Indexes content for Alexa answers and Amazon AI features. Respects robots.txt: Yes. Crackle recommendation: Allow — powers Alexa voice answers. Official docs →
cohere-ai — Operator: Cohere. User-Agent: cohere-ai. Purpose: Retrieval for Cohere Command models and enterprise RAG deployments. Respects robots.txt: Yes. Crackle recommendation: Allow — enterprise RAG systems increasingly cite via Cohere. Official docs →
Bytespider — Operator: ByteDance. User-Agent: Bytespider. Purpose: Training data for Doubao and TikTok AI features. Respects robots.txt: Yes. Crackle recommendation: Allow if you want APAC / TikTok-adjacent AI visibility. Official docs →
Training bots (GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, Meta-ExternalAgent, Bytespider). These fetch pages to train foundation models. Blocking them removes your brand from the corpus that becomes tomorrow's LLM knowledge.
Retrieval / indexing bots (OAI-SearchBot, PerplexityBot, Bingbot, Googlebot, claude-web, Amazonbot, cohere-ai). These build the real-time index the AI answer engine actually cites. Blocking them removes you from citation results directly.
User-initiated fetch bots (ChatGPT-User, Perplexity-User). Triggered by a live user prompt. Often treated as user browsers and may ignore robots.txt.
The default answer for almost every B2B tech brand: User-agent: * with Allow: /. Do not enumerate AI bots for blocking. Every AI bot is a distribution channel for your earned media and thought leadership.
The only exception is if a specific bot is causing measurable server load — in which case rate-limit at the CDN, don't block at robots.txt.
If you're on Cloudflare and enabled the AI Crawlers toggle at some point, disable it. Crackle PR treats robots.txt as the single source of truth to avoid conflicting signals.