Provider-documented reference

A source-backed map of AI fetchers, agents, and controls.

This directory separates automatic crawlers, user-triggered fetchers, and publisher control tokens using current provider documentation. It is a reference for interpreting candidate AI-agent traffic and crawler policy — not a claim that every listed identity has visited your storefront.

A user-agent string is not proof of identity, and provider behavior can change. Verify requests using the provider's current IP, DNS, or authentication guidance where available. Absence of a listed identity does not prove human traffic.

How to identify ChatGPT, Claude, and Perplexity traffic in your logs

Start with the request evidence you actually received, not the provider name alone. Read the claimed user-agent, referrer, provider token, or authenticated identity; distinguish an automatic crawler, a user-triggered fetcher, and a publisher control token; then verify the request against the provider's current official network, DNS, signature, or authentication guidance where available.

Retain the timestamp, path, method, status, referrer, relevant headers, evidence source, and verification result. If the identity cannot be verified, label it claimed / unverified. Never treat unknown traffic as human.

  1. Read the claimed user-agent, referrer, token, or authenticated identity evidence.
  2. Classify the request as an automatic crawler, user-triggered fetcher, or publisher control token.
  3. Verify against the provider's current official network, DNS, signature, or authentication guidance where published.
  4. Retain timestamp, path, method, status, referrer, relevant headers, evidence source, and verification result.
  5. Label unverifiable identity as claimed / unverified.
  6. Never infer that unknown traffic is human.

Documented entries

16

Operators represented

7

Sources checked

Aug 29, 2026

User-triggered fetchers and agents

5 entries
  • ChatGPT-User

    OpenAI

    user triggered fetcher

    User-triggered fetcher used for certain actions in ChatGPT and Custom GPTs; it is not an automatic web crawler.

    ChatGPT-User/1.0
    Robots behavior
    OpenAI says robots.txt rules may not apply because these requests are user-initiated.
    Identity evidence
    OpenAI publishes IP ranges for ChatGPT-User requests; the user-agent string remains spoofable.
    Primary provider source →
  • Claude-User

    Anthropic

    user triggered fetcher

    Retrieves web content in response to user-initiated Claude requests.

    Claude-User
    Robots behavior
    Anthropic says its bots honor robots.txt directives, including Claude-User.
    Identity evidence
    Anthropic's crawler guidance links a current source-IP list; do not trust the token alone.
    Primary provider source →
  • Perplexity-User

    Perplexity

    user triggered fetcher

    User-triggered fetcher that may visit a page to help answer a Perplexity user's question; it is not used for web crawling or foundation-model training.

    Perplexity-User/1.0
    Robots behavior
    Perplexity says this fetcher generally ignores robots.txt because the request was initiated by a user.
    Identity evidence
    Perplexity publishes IP ranges for Perplexity-User; combine the token and source IP rather than trusting the token alone.
    Primary provider source →
  • Google-Agent

    Google

    user triggered fetcher

    Used by agents hosted on Google infrastructure to navigate the web and perform actions upon user request.

    Google-Agent (browser-style user-agent string)
    Robots behavior
    Google says user-triggered fetchers generally ignore robots.txt rules because the fetch was requested by a user.
    Identity evidence
    Google publishes user-triggered-agent IP ranges and is experimenting with Web Bot Auth for the agent.bot.goog identity.
    Primary provider source →
  • Amzn-User

    Amazon

    user triggered fetcher

    Supports user actions such as fetching current web information to answer Alexa queries on a user's behalf; Amazon says it is not used for generative-AI model training.

    Amzn-User/0.1
    Robots behavior
    Amazon says because Amzn-User actions can be user-initiated, it may not follow all robots.txt directives.
    Identity evidence
    Amazon publishes IP addresses for Amzn-User; use source verification in addition to the token.
    Primary provider source →

Automatic search crawlers

5 entries
  • OAI-SearchBot

    OpenAI

    search crawler

    Automatic crawler used to surface websites in ChatGPT search results.

    OAI-SearchBot/1.4 (documented example; version may change)
    Robots behavior
    Managed with the OAI-SearchBot token in robots.txt.
    Identity evidence
    Match the token with OpenAI's published SearchBot IP ranges; a user-agent string alone is not proof.
    Primary provider source →
  • Claude-SearchBot

    Anthropic

    search crawler

    Navigates the web to improve the relevance and accuracy of Claude search responses.

    Claude-SearchBot
    Robots behavior
    Anthropic says its bots honor robots.txt directives.
    Identity evidence
    Anthropic's crawler guidance links a current source-IP list; do not trust the token alone.
    Primary provider source →
  • PerplexityBot

    Perplexity

    search crawler

    Automatic crawler designed to surface and link websites in Perplexity search results; Perplexity says it is not used to crawl content for foundation-model training.

    PerplexityBot/1.0
    Robots behavior
    Perplexity recommends managing PerplexityBot through robots.txt.
    Identity evidence
    Perplexity publishes IP ranges for PerplexityBot; combine the token and source IP rather than trusting the token alone.
    Primary provider source →
  • Applebot

    Apple

    search crawler

    Apple web crawler used for search experiences including Spotlight, Siri, and Safari; Apple says crawled data may also support foundation-model training and current-content context for AI outputs.

    Applebot/<version> within Apple's documented browser-style user-agent format
    Robots behavior
    Applebot respects standard robots.txt directives in general search crawls.
    Identity evidence
    Apple documents reverse-DNS verification under *.applebot.apple.com and publishes Applebot IP CIDR ranges.
    Primary provider source →
  • Amzn-SearchBot

    Amazon

    search crawler

    Amazon search crawler used to improve search experiences in Amazon products and services, including eligibility for experiences such as Alexa; Amazon says it is not used for generative-AI model training.

    Amzn-SearchBot/0.1
    Robots behavior
    Amazon documents robots.txt allow/disallow support for its crawlers.
    Identity evidence
    Amazon publishes IP addresses for Amzn-SearchBot; use source verification in addition to the token.
    Primary provider source →

Model-development crawlers

2 entries
  • GPTBot

    OpenAI

    model development crawler

    Automatic crawler for content that may be used to train OpenAI generative AI foundation models.

    GPTBot/1.4 (documented example; version may change)
    Robots behavior
    Managed separately with the GPTBot token in robots.txt.
    Identity evidence
    Match the token with OpenAI's published GPTBot IP ranges; a user-agent string alone is not proof.
    Primary provider source →
  • ClaudeBot

    Anthropic

    model development crawler

    Collects public web content that could potentially contribute to Anthropic model training.

    ClaudeBot
    Robots behavior
    Anthropic says its bots honor robots.txt directives.
    Identity evidence
    Anthropic's crawler guidance links a current source-IP list; do not trust the token alone.
    Primary provider source →

General provider crawlers

1 entry
  • Amazonbot

    Amazon

    general crawler

    Amazon crawler used to improve products and services; Amazon says collected content may also be used to train Amazon AI models.

    Amazonbot/0.1
    Robots behavior
    Amazon documents robots.txt allow/disallow support for its crawlers.
    Identity evidence
    Amazon publishes IP addresses for Amazonbot; use source verification in addition to the token.
    Primary provider source →

Publisher control tokens

2 entries
  • Google-Extended

    Google

    control token

    Publisher control for whether Google-crawled content may be used for future Gemini model training and specified grounding uses.

    robots.txt token; no separate HTTP request user-agent
    Robots behavior
    Google-Extended is a robots.txt product-control token and does not affect Google Search inclusion or ranking.
    Identity evidence
    Not a request identity. Do not look for Google-Extended as a standalone HTTP user-agent in traffic logs.
    Primary provider source →
  • Applebot-Extended

    Apple

    control token

    Publisher control for whether Applebot-crawled content may be used to train Apple's general-purpose foundation models.

    robots.txt control token; does not crawl webpages
    Robots behavior
    Applebot-Extended is configured in robots.txt but does not itself crawl webpages.
    Identity evidence
    Not a request identity. Traffic should be attributed to Applebot, not Applebot-Extended.
    Primary provider source →

Open web-corpus crawlers

1 entry
  • CCBot

    Common Crawl

    web corpus crawler

    Automated crawler that collects public web data for Common Crawl's open web-crawl repository.

    CCBot/2.0
    Robots behavior
    Common Crawl documents robots.txt support for CCBot.
    Identity evidence
    Common Crawl publishes dedicated IP ranges and reverse-DNS guidance; the project also warns that clients can falsely claim the CCBot user-agent.
    Primary provider source →

Three classes should not be collapsed into one "bot" label.

Automatic crawlers

Search and model-development crawlers operate automatically and are generally managed through provider-specific robots.txt policy.

User-triggered fetchers

These requests happen because a user or agent asked a provider to retrieve or act on a page. Robots behavior can differ materially from automatic crawlers.

Control tokens

Google-Extended and Applebot-Extended are publisher controls, not standalone HTTP crawler identities. They should not be treated as traffic labels.

What this directory establishes — and what it does not.

  • It records identities and controls that the named providers documented when this page was reviewed.
  • It does not establish that a request carrying one of these strings is legitimate; strings can be spoofed.
  • It does not establish that unrecognized traffic is human.
  • It does not establish what an external AI model concluded, preferred, recommended, compared, or decided.

Provider documentation checked 2026-08-29 · Re-check the linked primary source before operational enforcement

Cartograph is being built around evidence, not user-agent trust.

The planned evidence layer is intended to keep declared identity, verification evidence, observed storefront activity, and unknowns separate. Cartograph is not a bot blocker, WAF, or traffic-enforcement tool.