プロバイダーが文書化したリファレンス

AIフェッチャー、エージェント、およびコントロールのソースに基づいたマップ。

このディレクトリは、現在のプロバイダーのドキュメントを使用して、自動クローラー、ユーザーがトリガーするフェッチャー、およびパブリッシャー制御トークンを分離します。これは、候補のAIエージェントトラフィックとクローラーポリシーを解釈するための参照であり、リストされているすべてのIDがストアフロントを訪問したという主張ではありません。

ユーザーエージェント文字列は身元の証明ではなく、プロバイダーの動作は変更される可能性があります。プロバイダーの現在のIP、DNS、または利用可能な認証ガイダンスを使用してリクエストを確認してください。リストに記載された身元がないことは、人間によるトラフィックであることを証明するものではありません。

文書化されたエントリ

16

担当オペレーター

7

確認済みのソース

Aug 29, 2026

ユーザーがトリガーするフェッチャーとエージェント

5 entries
  • ChatGPT-User

    OpenAI

    user triggered fetcher

    User-triggered fetcher used for certain actions in ChatGPT and Custom GPTs; it is not an automatic web crawler.

    ChatGPT-User/1.0
    ロボットの動作
    OpenAI says robots.txt rules may not apply because these requests are user-initiated.
    IDエビデンス
    OpenAI publishes IP ranges for ChatGPT-User requests; the user-agent string remains spoofable.
    主要プロバイダーソース →
  • Claude-User

    Anthropic

    user triggered fetcher

    Retrieves web content in response to user-initiated Claude requests.

    Claude-User
    ロボットの動作
    Anthropic says its bots honor robots.txt directives, including Claude-User.
    IDエビデンス
    Anthropic's crawler guidance links a current source-IP list; do not trust the token alone.
    主要プロバイダーソース →
  • Perplexity-User

    Perplexity

    user triggered fetcher

    User-triggered fetcher that may visit a page to help answer a Perplexity user's question; it is not used for web crawling or foundation-model training.

    Perplexity-User/1.0
    ロボットの動作
    Perplexity says this fetcher generally ignores robots.txt because the request was initiated by a user.
    IDエビデンス
    Perplexity publishes IP ranges for Perplexity-User; combine the token and source IP rather than trusting the token alone.
    主要プロバイダーソース →
  • Google-Agent

    Google

    user triggered fetcher

    Used by agents hosted on Google infrastructure to navigate the web and perform actions upon user request.

    Google-Agent (browser-style user-agent string)
    ロボットの動作
    Google says user-triggered fetchers generally ignore robots.txt rules because the fetch was requested by a user.
    IDエビデンス
    Google publishes user-triggered-agent IP ranges and is experimenting with Web Bot Auth for the agent.bot.goog identity.
    主要プロバイダーソース →
  • Amzn-User

    Amazon

    user triggered fetcher

    Supports user actions such as fetching current web information to answer Alexa queries on a user's behalf; Amazon says it is not used for generative-AI model training.

    Amzn-User/0.1
    ロボットの動作
    Amazon says because Amzn-User actions can be user-initiated, it may not follow all robots.txt directives.
    IDエビデンス
    Amazon publishes IP addresses for Amzn-User; use source verification in addition to the token.
    主要プロバイダーソース →

自動検索クローラー

5 entries
  • OAI-SearchBot

    OpenAI

    search crawler

    Automatic crawler used to surface websites in ChatGPT search results.

    OAI-SearchBot/1.4 (documented example; version may change)
    ロボットの動作
    Managed with the OAI-SearchBot token in robots.txt.
    IDエビデンス
    Match the token with OpenAI's published SearchBot IP ranges; a user-agent string alone is not proof.
    主要プロバイダーソース →
  • Claude-SearchBot

    Anthropic

    search crawler

    Navigates the web to improve the relevance and accuracy of Claude search responses.

    Claude-SearchBot
    ロボットの動作
    Anthropic says its bots honor robots.txt directives.
    IDエビデンス
    Anthropic's crawler guidance links a current source-IP list; do not trust the token alone.
    主要プロバイダーソース →
  • PerplexityBot

    Perplexity

    search crawler

    Automatic crawler designed to surface and link websites in Perplexity search results; Perplexity says it is not used to crawl content for foundation-model training.

    PerplexityBot/1.0
    ロボットの動作
    Perplexity recommends managing PerplexityBot through robots.txt.
    IDエビデンス
    Perplexity publishes IP ranges for PerplexityBot; combine the token and source IP rather than trusting the token alone.
    主要プロバイダーソース →
  • Applebot

    Apple

    search crawler

    Apple web crawler used for search experiences including Spotlight, Siri, and Safari; Apple says crawled data may also support foundation-model training and current-content context for AI outputs.

    Applebot/<version> within Apple's documented browser-style user-agent format
    ロボットの動作
    Applebot respects standard robots.txt directives in general search crawls.
    IDエビデンス
    Apple documents reverse-DNS verification under *.applebot.apple.com and publishes Applebot IP CIDR ranges.
    主要プロバイダーソース →
  • Amzn-SearchBot

    Amazon

    search crawler

    Amazon search crawler used to improve search experiences in Amazon products and services, including eligibility for experiences such as Alexa; Amazon says it is not used for generative-AI model training.

    Amzn-SearchBot/0.1
    ロボットの動作
    Amazon documents robots.txt allow/disallow support for its crawlers.
    IDエビデンス
    Amazon publishes IP addresses for Amzn-SearchBot; use source verification in addition to the token.
    主要プロバイダーソース →

モデル開発クローラー

2 entries
  • GPTBot

    OpenAI

    model development crawler

    Automatic crawler for content that may be used to train OpenAI generative AI foundation models.

    GPTBot/1.4 (documented example; version may change)
    ロボットの動作
    Managed separately with the GPTBot token in robots.txt.
    IDエビデンス
    Match the token with OpenAI's published GPTBot IP ranges; a user-agent string alone is not proof.
    主要プロバイダーソース →
  • ClaudeBot

    Anthropic

    model development crawler

    Collects public web content that could potentially contribute to Anthropic model training.

    ClaudeBot
    ロボットの動作
    Anthropic says its bots honor robots.txt directives.
    IDエビデンス
    Anthropic's crawler guidance links a current source-IP list; do not trust the token alone.
    主要プロバイダーソース →

一般的なプロバイダーのクローラー

1 entry
  • Amazonbot

    Amazon

    general crawler

    Amazon crawler used to improve products and services; Amazon says collected content may also be used to train Amazon AI models.

    Amazonbot/0.1
    ロボットの動作
    Amazon documents robots.txt allow/disallow support for its crawlers.
    IDエビデンス
    Amazon publishes IP addresses for Amazonbot; use source verification in addition to the token.
    主要プロバイダーソース →

パブリッシャーコントロールトークン

2 entries
  • Google-Extended

    Google

    control token

    Publisher control for whether Google-crawled content may be used for future Gemini model training and specified grounding uses.

    robots.txt token; no separate HTTP request user-agent
    ロボットの動作
    Google-Extended is a robots.txt product-control token and does not affect Google Search inclusion or ranking.
    IDエビデンス
    Not a request identity. Do not look for Google-Extended as a standalone HTTP user-agent in traffic logs.
    主要プロバイダーソース →
  • Applebot-Extended

    Apple

    control token

    Publisher control for whether Applebot-crawled content may be used to train Apple's general-purpose foundation models.

    robots.txt control token; does not crawl webpages
    ロボットの動作
    Applebot-Extended is configured in robots.txt but does not itself crawl webpages.
    IDエビデンス
    Not a request identity. Traffic should be attributed to Applebot, not Applebot-Extended.
    主要プロバイダーソース →

オープンウェブコーパスクローラー

1 entry
  • CCBot

    Common Crawl

    web corpus crawler

    Automated crawler that collects public web data for Common Crawl's open web-crawl repository.

    CCBot/2.0
    ロボットの動作
    Common Crawl documents robots.txt support for CCBot.
    IDエビデンス
    Common Crawl publishes dedicated IP ranges and reverse-DNS guidance; the project also warns that clients can falsely claim the CCBot user-agent.
    主要プロバイダーソース →

3つのクラスを1つの「ボット」ラベルにまとめるべきではありません。

自動クローラー

Search and model-development crawlers operate automatically and are generally managed through provider-specific robots.txt policy.

ユーザーがトリガーするフェッチャー

These requests happen because a user or agent asked a provider to retrieve or act on a page. Robots behavior can differ materially from automatic crawlers.

制御トークン

Google-Extended and Applebot-Extended are publisher controls, not standalone HTTP crawler identities. They should not be treated as traffic labels.

このディレクトリが確立しているもの — そして確立していないもの。

  • このページがレビューされた時点で、指定されたプロバイダーが文書化したIDと制御を記録します。
  • これらの文字列のいずれかを運ぶリクエストが正当であることを確立するものではありません。文字列は偽装される可能性があります。
  • 認識されていないトラフィックが人間であることを確立するものではありません。
  • 外部AIモデルが何を結論付け、好み、推奨し、比較し、決定したかを確立するものではありません。

プロバイダーのドキュメントを確認済み 2026-08-29 · 運用の実施前にリンクされた主要なソースを再確認してください

Cartographは、ユーザーエージェントの信頼ではなく、証拠に基づいて構築されています。

計画されたエビデンスレイヤーは、宣言されたID、検証エビデンス、観測されたストアフロントアクティビティ、および不明なものを分離することを目的としています。Cartographは、ボットブロッカー、WAF、またはトラフィック強制ツールではありません。