GPTBot
OpenAI
Foundation-model training crawler. Distinct from ChatGPT search and user fetches.
Registry
25 agents · GET /v1/agents
GPTBot
OpenAI
Foundation-model training crawler. Distinct from ChatGPT search and user fetches.
ChatGPT-User
OpenAI
On-demand fetch when a person asks ChatGPT to read a page.
OAI-SearchBot
OpenAI
Search indexing for ChatGPT search. Not a training crawler.
ClaudeBot
Anthropic
Training crawler for Claude models.
Claude-User
Anthropic
User-initiated fetch from Claude.
Claude-SearchBot
Anthropic
Improves Claude search result quality.
Google-Extended
robots.txt control token, not a separate User-Agent. Governs Gemini training and grounding. Does not affect Google Search ranking.
Googlebot
Web search crawler. Not a Gemini training signal.
Bingbot
Microsoft
Bing and Copilot retrieval depend on Bing's index.
Applebot
Apple
Powers Spotlight, Siri suggestions, and Safari. Not Apple Intelligence training.
Applebot-Extended
Apple
Control token for Apple Intelligence training. Separate from Applebot search.
PerplexityBot
Perplexity
Builds Perplexity's search index.
Perplexity-User
Perplexity
User-initiated fetch from a Perplexity answer.
Amazonbot
Amazon
Supports Alexa answers and Amazon services.
Meta-ExternalAgent
Meta
AI training and ranking for Meta AI products.
FacebookBot
Meta
Documented as a training crawler.
CCBot
Common Crawl
Open web corpus. Widely reused as training data by other labs.
Bytespider
ByteDance
TikTok / ByteDance model training crawler.
cohere-ai
Cohere
User-initiated retrieval for Cohere answers.
cohere-training-data-crawler
Cohere
Training-data crawler. Separate from cohere-ai retrieval.
Diffbot
Diffbot
Knowledge-graph extraction used by downstream AI products.
DuckAssistBot
DuckDuckGo
Fetches sources for Duck.ai answers.
DeepSeekBot
DeepSeek
Training and product-improvement crawler.
GrokBot
xAI
robots.txt token used to opt in or out of xAI training crawls.
YouBot
You.com
You.com search and assistant retrieval.