Um mapa de fetchers, agentes e controles de IA com respaldo de origem.
Este diretório separa rastreadores automáticos, buscadores acionados por usuários e tokens de controle de editor usando a documentação atual do provedor. É uma referência para interpretar o tráfego de agentes de IA candidatos e a política de rastreadores — não uma afirmação de que toda identidade listada visitou sua vitrine.
Um user-agent string não é prova de identidade, e o comportamento do provedor pode mudar. Verifique as solicitações usando o IP atual do provedor, DNS ou orientação de autenticação onde disponível. A ausência de uma identidade listada não prova tráfego humano.
Entradas documentadas
16
Operadores representados
7
Fontes verificadas
Aug 29, 2026
Agentes e fetchers acionados pelo usuário
5 entries- user triggered fetcher
ChatGPT-User
OpenAI
User-triggered fetcher used for certain actions in ChatGPT and Custom GPTs; it is not an automatic web crawler.
ChatGPT-User/1.0- Comportamento de robôs
- OpenAI says robots.txt rules may not apply because these requests are user-initiated.
- Evidência de identidade
- OpenAI publishes IP ranges for ChatGPT-User requests; the user-agent string remains spoofable.
- user triggered fetcher
Claude-User
Anthropic
Retrieves web content in response to user-initiated Claude requests.
Claude-User- Comportamento de robôs
- Anthropic says its bots honor robots.txt directives, including Claude-User.
- Evidência de identidade
- Anthropic's crawler guidance links a current source-IP list; do not trust the token alone.
- user triggered fetcher
Perplexity-User
Perplexity
User-triggered fetcher that may visit a page to help answer a Perplexity user's question; it is not used for web crawling or foundation-model training.
Perplexity-User/1.0- Comportamento de robôs
- Perplexity says this fetcher generally ignores robots.txt because the request was initiated by a user.
- Evidência de identidade
- Perplexity publishes IP ranges for Perplexity-User; combine the token and source IP rather than trusting the token alone.
- user triggered fetcher
Google-Agent
Google
Used by agents hosted on Google infrastructure to navigate the web and perform actions upon user request.
Google-Agent (browser-style user-agent string)- Comportamento de robôs
- Google says user-triggered fetchers generally ignore robots.txt rules because the fetch was requested by a user.
- Evidência de identidade
- Google publishes user-triggered-agent IP ranges and is experimenting with Web Bot Auth for the agent.bot.goog identity.
- user triggered fetcher
Amzn-User
Amazon
Supports user actions such as fetching current web information to answer Alexa queries on a user's behalf; Amazon says it is not used for generative-AI model training.
Amzn-User/0.1- Comportamento de robôs
- Amazon says because Amzn-User actions can be user-initiated, it may not follow all robots.txt directives.
- Evidência de identidade
- Amazon publishes IP addresses for Amzn-User; use source verification in addition to the token.
Rastreadores de busca automáticos
5 entries- search crawler
OAI-SearchBot
OpenAI
Automatic crawler used to surface websites in ChatGPT search results.
OAI-SearchBot/1.4 (documented example; version may change)- Comportamento de robôs
- Managed with the OAI-SearchBot token in robots.txt.
- Evidência de identidade
- Match the token with OpenAI's published SearchBot IP ranges; a user-agent string alone is not proof.
- search crawler
Claude-SearchBot
Anthropic
Navigates the web to improve the relevance and accuracy of Claude search responses.
Claude-SearchBot- Comportamento de robôs
- Anthropic says its bots honor robots.txt directives.
- Evidência de identidade
- Anthropic's crawler guidance links a current source-IP list; do not trust the token alone.
- search crawler
PerplexityBot
Perplexity
Automatic crawler designed to surface and link websites in Perplexity search results; Perplexity says it is not used to crawl content for foundation-model training.
PerplexityBot/1.0- Comportamento de robôs
- Perplexity recommends managing PerplexityBot through robots.txt.
- Evidência de identidade
- Perplexity publishes IP ranges for PerplexityBot; combine the token and source IP rather than trusting the token alone.
- search crawler
Applebot
Apple
Apple web crawler used for search experiences including Spotlight, Siri, and Safari; Apple says crawled data may also support foundation-model training and current-content context for AI outputs.
Applebot/<version> within Apple's documented browser-style user-agent format- Comportamento de robôs
- Applebot respects standard robots.txt directives in general search crawls.
- Evidência de identidade
- Apple documents reverse-DNS verification under *.applebot.apple.com and publishes Applebot IP CIDR ranges.
- search crawler
Amzn-SearchBot
Amazon
Amazon search crawler used to improve search experiences in Amazon products and services, including eligibility for experiences such as Alexa; Amazon says it is not used for generative-AI model training.
Amzn-SearchBot/0.1- Comportamento de robôs
- Amazon documents robots.txt allow/disallow support for its crawlers.
- Evidência de identidade
- Amazon publishes IP addresses for Amzn-SearchBot; use source verification in addition to the token.
Rastreadores de desenvolvimento de modelo
2 entries- model development crawler
GPTBot
OpenAI
Automatic crawler for content that may be used to train OpenAI generative AI foundation models.
GPTBot/1.4 (documented example; version may change)- Comportamento de robôs
- Managed separately with the GPTBot token in robots.txt.
- Evidência de identidade
- Match the token with OpenAI's published GPTBot IP ranges; a user-agent string alone is not proof.
- model development crawler
ClaudeBot
Anthropic
Collects public web content that could potentially contribute to Anthropic model training.
ClaudeBot- Comportamento de robôs
- Anthropic says its bots honor robots.txt directives.
- Evidência de identidade
- Anthropic's crawler guidance links a current source-IP list; do not trust the token alone.
Rastreadores gerais de provedores
1 entry- general crawler
Amazonbot
Amazon
Amazon crawler used to improve products and services; Amazon says collected content may also be used to train Amazon AI models.
Amazonbot/0.1- Comportamento de robôs
- Amazon documents robots.txt allow/disallow support for its crawlers.
- Evidência de identidade
- Amazon publishes IP addresses for Amazonbot; use source verification in addition to the token.
Tokens de controle do editor
2 entries- control token
Google-Extended
Google
Publisher control for whether Google-crawled content may be used for future Gemini model training and specified grounding uses.
robots.txt token; no separate HTTP request user-agent- Comportamento de robôs
- Google-Extended is a robots.txt product-control token and does not affect Google Search inclusion or ranking.
- Evidência de identidade
- Not a request identity. Do not look for Google-Extended as a standalone HTTP user-agent in traffic logs.
- control token
Applebot-Extended
Apple
Publisher control for whether Applebot-crawled content may be used to train Apple's general-purpose foundation models.
robots.txt control token; does not crawl webpages- Comportamento de robôs
- Applebot-Extended is configured in robots.txt but does not itself crawl webpages.
- Evidência de identidade
- Not a request identity. Traffic should be attributed to Applebot, not Applebot-Extended.
Rastreadores de corpus web aberto
1 entry- web corpus crawler
CCBot
Common Crawl
Automated crawler that collects public web data for Common Crawl's open web-crawl repository.
CCBot/2.0- Comportamento de robôs
- Common Crawl documents robots.txt support for CCBot.
- Evidência de identidade
- Common Crawl publishes dedicated IP ranges and reverse-DNS guidance; the project also warns that clients can falsely claim the CCBot user-agent.
Três classes não devem ser colapsadas em um único rótulo "bot".
Rastreadores automáticos
Search and model-development crawlers operate automatically and are generally managed through provider-specific robots.txt policy.
Fetchers acionados pelo usuário
These requests happen because a user or agent asked a provider to retrieve or act on a page. Robots behavior can differ materially from automatic crawlers.
Tokens de controle
Google-Extended and Applebot-Extended are publisher controls, not standalone HTTP crawler identities. They should not be treated as traffic labels.
O que este diretório estabelece — e o que não estabelece.
- Ele registra identidades e controles que os provedores nomeados documentaram quando esta página foi revisada.
- Não estabelece que uma solicitação contendo uma dessas strings seja legítima; as strings podem ser falsificadas.
- Não estabelece que tráfego não reconhecido seja humano.
- Não estabelece o que um modelo de IA externo concluiu, preferiu, recomendou, comparou ou decidiu.
Documentação do provedor verificada 2026-08-29 · Re-verifique a fonte primária vinculada antes da aplicação operacional
Cartograph está sendo construído em torno de evidências, não na confiança do user-agent.
A camada de evidências planejada visa manter separadas a identidade declarada, as evidências de verificação, a atividade observada da vitrine e os itens desconhecidos. A Cartograph não é um bloqueador de bots, WAF ou ferramenta de aplicação de tráfego.