AIフェッチャー、エージェント、およびコントロールのソースに基づいたマップ。
このディレクトリは、現在のプロバイダーのドキュメントを使用して、自動クローラー、ユーザーがトリガーするフェッチャー、およびパブリッシャー制御トークンを分離します。これは、候補のAIエージェントトラフィックとクローラーポリシーを解釈するための参照であり、リストされているすべてのIDがストアフロントを訪問したという主張ではありません。
ユーザーエージェント文字列は身元の証明ではなく、プロバイダーの動作は変更される可能性があります。プロバイダーの現在のIP、DNS、または利用可能な認証ガイダンスを使用してリクエストを確認してください。リストに記載された身元がないことは、人間によるトラフィックであることを証明するものではありません。
文書化されたエントリ
16
担当オペレーター
7
確認済みのソース
Aug 29, 2026
ユーザーがトリガーするフェッチャーとエージェント
5 entries- user triggered fetcher
ChatGPT-User
OpenAI
User-triggered fetcher used for certain actions in ChatGPT and Custom GPTs; it is not an automatic web crawler.
ChatGPT-User/1.0- ロボットの動作
- OpenAI says robots.txt rules may not apply because these requests are user-initiated.
- IDエビデンス
- OpenAI publishes IP ranges for ChatGPT-User requests; the user-agent string remains spoofable.
- user triggered fetcher
Claude-User
Anthropic
Retrieves web content in response to user-initiated Claude requests.
Claude-User- ロボットの動作
- Anthropic says its bots honor robots.txt directives, including Claude-User.
- IDエビデンス
- Anthropic's crawler guidance links a current source-IP list; do not trust the token alone.
- user triggered fetcher
Perplexity-User
Perplexity
User-triggered fetcher that may visit a page to help answer a Perplexity user's question; it is not used for web crawling or foundation-model training.
Perplexity-User/1.0- ロボットの動作
- Perplexity says this fetcher generally ignores robots.txt because the request was initiated by a user.
- IDエビデンス
- Perplexity publishes IP ranges for Perplexity-User; combine the token and source IP rather than trusting the token alone.
- user triggered fetcher
Google-Agent
Google
Used by agents hosted on Google infrastructure to navigate the web and perform actions upon user request.
Google-Agent (browser-style user-agent string)- ロボットの動作
- Google says user-triggered fetchers generally ignore robots.txt rules because the fetch was requested by a user.
- IDエビデンス
- Google publishes user-triggered-agent IP ranges and is experimenting with Web Bot Auth for the agent.bot.goog identity.
- user triggered fetcher
Amzn-User
Amazon
Supports user actions such as fetching current web information to answer Alexa queries on a user's behalf; Amazon says it is not used for generative-AI model training.
Amzn-User/0.1- ロボットの動作
- Amazon says because Amzn-User actions can be user-initiated, it may not follow all robots.txt directives.
- IDエビデンス
- Amazon publishes IP addresses for Amzn-User; use source verification in addition to the token.
自動検索クローラー
5 entries- search crawler
OAI-SearchBot
OpenAI
Automatic crawler used to surface websites in ChatGPT search results.
OAI-SearchBot/1.4 (documented example; version may change)- ロボットの動作
- Managed with the OAI-SearchBot token in robots.txt.
- IDエビデンス
- Match the token with OpenAI's published SearchBot IP ranges; a user-agent string alone is not proof.
- search crawler
Claude-SearchBot
Anthropic
Navigates the web to improve the relevance and accuracy of Claude search responses.
Claude-SearchBot- ロボットの動作
- Anthropic says its bots honor robots.txt directives.
- IDエビデンス
- Anthropic's crawler guidance links a current source-IP list; do not trust the token alone.
- search crawler
PerplexityBot
Perplexity
Automatic crawler designed to surface and link websites in Perplexity search results; Perplexity says it is not used to crawl content for foundation-model training.
PerplexityBot/1.0- ロボットの動作
- Perplexity recommends managing PerplexityBot through robots.txt.
- IDエビデンス
- Perplexity publishes IP ranges for PerplexityBot; combine the token and source IP rather than trusting the token alone.
- search crawler
Applebot
Apple
Apple web crawler used for search experiences including Spotlight, Siri, and Safari; Apple says crawled data may also support foundation-model training and current-content context for AI outputs.
Applebot/<version> within Apple's documented browser-style user-agent format- ロボットの動作
- Applebot respects standard robots.txt directives in general search crawls.
- IDエビデンス
- Apple documents reverse-DNS verification under *.applebot.apple.com and publishes Applebot IP CIDR ranges.
- search crawler
Amzn-SearchBot
Amazon
Amazon search crawler used to improve search experiences in Amazon products and services, including eligibility for experiences such as Alexa; Amazon says it is not used for generative-AI model training.
Amzn-SearchBot/0.1- ロボットの動作
- Amazon documents robots.txt allow/disallow support for its crawlers.
- IDエビデンス
- Amazon publishes IP addresses for Amzn-SearchBot; use source verification in addition to the token.
モデル開発クローラー
2 entries- model development crawler
GPTBot
OpenAI
Automatic crawler for content that may be used to train OpenAI generative AI foundation models.
GPTBot/1.4 (documented example; version may change)- ロボットの動作
- Managed separately with the GPTBot token in robots.txt.
- IDエビデンス
- Match the token with OpenAI's published GPTBot IP ranges; a user-agent string alone is not proof.
- model development crawler
ClaudeBot
Anthropic
Collects public web content that could potentially contribute to Anthropic model training.
ClaudeBot- ロボットの動作
- Anthropic says its bots honor robots.txt directives.
- IDエビデンス
- Anthropic's crawler guidance links a current source-IP list; do not trust the token alone.
一般的なプロバイダーのクローラー
1 entry- general crawler
Amazonbot
Amazon
Amazon crawler used to improve products and services; Amazon says collected content may also be used to train Amazon AI models.
Amazonbot/0.1- ロボットの動作
- Amazon documents robots.txt allow/disallow support for its crawlers.
- IDエビデンス
- Amazon publishes IP addresses for Amazonbot; use source verification in addition to the token.
パブリッシャーコントロールトークン
2 entries- control token
Google-Extended
Google
Publisher control for whether Google-crawled content may be used for future Gemini model training and specified grounding uses.
robots.txt token; no separate HTTP request user-agent- ロボットの動作
- Google-Extended is a robots.txt product-control token and does not affect Google Search inclusion or ranking.
- IDエビデンス
- Not a request identity. Do not look for Google-Extended as a standalone HTTP user-agent in traffic logs.
- control token
Applebot-Extended
Apple
Publisher control for whether Applebot-crawled content may be used to train Apple's general-purpose foundation models.
robots.txt control token; does not crawl webpages- ロボットの動作
- Applebot-Extended is configured in robots.txt but does not itself crawl webpages.
- IDエビデンス
- Not a request identity. Traffic should be attributed to Applebot, not Applebot-Extended.
オープンウェブコーパスクローラー
1 entry- web corpus crawler
CCBot
Common Crawl
Automated crawler that collects public web data for Common Crawl's open web-crawl repository.
CCBot/2.0- ロボットの動作
- Common Crawl documents robots.txt support for CCBot.
- IDエビデンス
- Common Crawl publishes dedicated IP ranges and reverse-DNS guidance; the project also warns that clients can falsely claim the CCBot user-agent.
3つのクラスを1つの「ボット」ラベルにまとめるべきではありません。
自動クローラー
Search and model-development crawlers operate automatically and are generally managed through provider-specific robots.txt policy.
ユーザーがトリガーするフェッチャー
These requests happen because a user or agent asked a provider to retrieve or act on a page. Robots behavior can differ materially from automatic crawlers.
制御トークン
Google-Extended and Applebot-Extended are publisher controls, not standalone HTTP crawler identities. They should not be treated as traffic labels.
このディレクトリが確立しているもの — そして確立していないもの。
- このページがレビューされた時点で、指定されたプロバイダーが文書化したIDと制御を記録します。
- これらの文字列のいずれかを運ぶリクエストが正当であることを確立するものではありません。文字列は偽装される可能性があります。
- 認識されていないトラフィックが人間であることを確立するものではありません。
- 外部AIモデルが何を結論付け、好み、推奨し、比較し、決定したかを確立するものではありません。
プロバイダーのドキュメントを確認済み 2026-08-29 · 運用の実施前にリンクされた主要なソースを再確認してください
Cartographは、ユーザーエージェントの信頼ではなく、証拠に基づいて構築されています。
計画されたエビデンスレイヤーは、宣言されたID、検証エビデンス、観測されたストアフロントアクティビティ、および不明なものを分離することを目的としています。Cartographは、ボットブロッカー、WAF、またはトラフィック強制ツールではありません。