Registry

Known agents

The vocabulary robots.txt actually understands. Control tokens are not HTTP User-Agents. User fetches happen because a person asked — they are not training crawls.
25 tokens

25 agents · GET /v1/agents

GPTBot

OpenAI

crawlertrain

Foundation-model training crawler. Distinct from ChatGPT search and user fetches.

ChatGPT-User

OpenAI

user-fetchretrieveact

On-demand fetch when a person asks ChatGPT to read a page.

OAI-SearchBot

OpenAI

crawlersearchretrieve

Search indexing for ChatGPT search. Not a training crawler.

ClaudeBot

Anthropic

crawlertrain

Training crawler for Claude models.

Claude-User

Anthropic

user-fetchretrieveact

User-initiated fetch from Claude.

Claude-SearchBot

Anthropic

crawlersearchretrieve

Improves Claude search result quality.

Google-Extended

Google

control-tokentrain

robots.txt control token, not a separate User-Agent. Governs Gemini training and grounding. Does not affect Google Search ranking.

Googlebot

Google

crawlersearch

Web search crawler. Not a Gemini training signal.

Bingbot

Microsoft

crawlersearchretrieve

Bing and Copilot retrieval depend on Bing's index.

Applebot

Apple

crawlersearch

Powers Spotlight, Siri suggestions, and Safari. Not Apple Intelligence training.

Applebot-Extended

Apple

control-tokentrain

Control token for Apple Intelligence training. Separate from Applebot search.

PerplexityBot

Perplexity

crawlersearchretrieve

Builds Perplexity's search index.

Perplexity-User

Perplexity

user-fetchretrieveact

User-initiated fetch from a Perplexity answer.

Amazonbot

Amazon

crawlerretrievesearch

Supports Alexa answers and Amazon services.

Meta-ExternalAgent

Meta

crawlertrainretrieve

AI training and ranking for Meta AI products.

FacebookBot

Meta

crawlertrain

Documented as a training crawler.

CCBot

Common Crawl

crawlertrain

Open web corpus. Widely reused as training data by other labs.

Bytespider

ByteDance

crawlertrain

TikTok / ByteDance model training crawler.

cohere-ai

Cohere

user-fetchretrieve

User-initiated retrieval for Cohere answers.

cohere-training-data-crawler

Cohere

crawlertrain

Training-data crawler. Separate from cohere-ai retrieval.

Diffbot

Diffbot

crawlerretrievetrain

Knowledge-graph extraction used by downstream AI products.

DuckAssistBot

DuckDuckGo

crawlerretrieve

Fetches sources for Duck.ai answers.

DeepSeekBot

DeepSeek

crawlertrain

Training and product-improvement crawler.

GrokBot

xAI

crawlertrain

robots.txt token used to opt in or out of xAI training crawls.

YouBot

You.com

crawlersearchretrieve

You.com search and assistant retrieval.