提供商记录的参考

一份由来源支持的 AI 抓取器、代理和控制的地图。

此目录使用当前提供商文档将自动爬虫、用户触发的抓取工具和发布者控制令牌分开。它是解释候选 AI 代理流量和爬虫策略的参考——而不是声称每个列出的身份都访问过您的店面。

用户代理字符串并非身份证明,且提供商行为可能发生变化。在可能的情况下,使用提供商当前的 IP、DNS 或认证指南验证请求。未列出的身份不代表是人类流量。

已记录条目

16

已代表的运营商

7

已检查来源

Aug 29, 2026

用户触发的获取器和代理

5 entries
  • ChatGPT-User

    OpenAI

    user triggered fetcher

    User-triggered fetcher used for certain actions in ChatGPT and Custom GPTs; it is not an automatic web crawler.

    ChatGPT-User/1.0
    机器人行为
    OpenAI says robots.txt rules may not apply because these requests are user-initiated.
    身份证据
    OpenAI publishes IP ranges for ChatGPT-User requests; the user-agent string remains spoofable.
    主要提供商来源 →
  • Claude-User

    Anthropic

    user triggered fetcher

    Retrieves web content in response to user-initiated Claude requests.

    Claude-User
    机器人行为
    Anthropic says its bots honor robots.txt directives, including Claude-User.
    身份证据
    Anthropic's crawler guidance links a current source-IP list; do not trust the token alone.
    主要提供商来源 →
  • Perplexity-User

    Perplexity

    user triggered fetcher

    User-triggered fetcher that may visit a page to help answer a Perplexity user's question; it is not used for web crawling or foundation-model training.

    Perplexity-User/1.0
    机器人行为
    Perplexity says this fetcher generally ignores robots.txt because the request was initiated by a user.
    身份证据
    Perplexity publishes IP ranges for Perplexity-User; combine the token and source IP rather than trusting the token alone.
    主要提供商来源 →
  • Google-Agent

    Google

    user triggered fetcher

    Used by agents hosted on Google infrastructure to navigate the web and perform actions upon user request.

    Google-Agent (browser-style user-agent string)
    机器人行为
    Google says user-triggered fetchers generally ignore robots.txt rules because the fetch was requested by a user.
    身份证据
    Google publishes user-triggered-agent IP ranges and is experimenting with Web Bot Auth for the agent.bot.goog identity.
    主要提供商来源 →
  • Amzn-User

    Amazon

    user triggered fetcher

    Supports user actions such as fetching current web information to answer Alexa queries on a user's behalf; Amazon says it is not used for generative-AI model training.

    Amzn-User/0.1
    机器人行为
    Amazon says because Amzn-User actions can be user-initiated, it may not follow all robots.txt directives.
    身份证据
    Amazon publishes IP addresses for Amzn-User; use source verification in addition to the token.
    主要提供商来源 →

自动搜索爬虫

5 entries
  • OAI-SearchBot

    OpenAI

    search crawler

    Automatic crawler used to surface websites in ChatGPT search results.

    OAI-SearchBot/1.4 (documented example; version may change)
    机器人行为
    Managed with the OAI-SearchBot token in robots.txt.
    身份证据
    Match the token with OpenAI's published SearchBot IP ranges; a user-agent string alone is not proof.
    主要提供商来源 →
  • Claude-SearchBot

    Anthropic

    search crawler

    Navigates the web to improve the relevance and accuracy of Claude search responses.

    Claude-SearchBot
    机器人行为
    Anthropic says its bots honor robots.txt directives.
    身份证据
    Anthropic's crawler guidance links a current source-IP list; do not trust the token alone.
    主要提供商来源 →
  • PerplexityBot

    Perplexity

    search crawler

    Automatic crawler designed to surface and link websites in Perplexity search results; Perplexity says it is not used to crawl content for foundation-model training.

    PerplexityBot/1.0
    机器人行为
    Perplexity recommends managing PerplexityBot through robots.txt.
    身份证据
    Perplexity publishes IP ranges for PerplexityBot; combine the token and source IP rather than trusting the token alone.
    主要提供商来源 →
  • Applebot

    Apple

    search crawler

    Apple web crawler used for search experiences including Spotlight, Siri, and Safari; Apple says crawled data may also support foundation-model training and current-content context for AI outputs.

    Applebot/<version> within Apple's documented browser-style user-agent format
    机器人行为
    Applebot respects standard robots.txt directives in general search crawls.
    身份证据
    Apple documents reverse-DNS verification under *.applebot.apple.com and publishes Applebot IP CIDR ranges.
    主要提供商来源 →
  • Amzn-SearchBot

    Amazon

    search crawler

    Amazon search crawler used to improve search experiences in Amazon products and services, including eligibility for experiences such as Alexa; Amazon says it is not used for generative-AI model training.

    Amzn-SearchBot/0.1
    机器人行为
    Amazon documents robots.txt allow/disallow support for its crawlers.
    身份证据
    Amazon publishes IP addresses for Amzn-SearchBot; use source verification in addition to the token.
    主要提供商来源 →

模型开发爬虫

2 entries
  • GPTBot

    OpenAI

    model development crawler

    Automatic crawler for content that may be used to train OpenAI generative AI foundation models.

    GPTBot/1.4 (documented example; version may change)
    机器人行为
    Managed separately with the GPTBot token in robots.txt.
    身份证据
    Match the token with OpenAI's published GPTBot IP ranges; a user-agent string alone is not proof.
    主要提供商来源 →
  • ClaudeBot

    Anthropic

    model development crawler

    Collects public web content that could potentially contribute to Anthropic model training.

    ClaudeBot
    机器人行为
    Anthropic says its bots honor robots.txt directives.
    身份证据
    Anthropic's crawler guidance links a current source-IP list; do not trust the token alone.
    主要提供商来源 →

通用提供商爬虫

1 entry
  • Amazonbot

    Amazon

    general crawler

    Amazon crawler used to improve products and services; Amazon says collected content may also be used to train Amazon AI models.

    Amazonbot/0.1
    机器人行为
    Amazon documents robots.txt allow/disallow support for its crawlers.
    身份证据
    Amazon publishes IP addresses for Amazonbot; use source verification in addition to the token.
    主要提供商来源 →

发布者控制令牌

2 entries
  • Google-Extended

    Google

    control token

    Publisher control for whether Google-crawled content may be used for future Gemini model training and specified grounding uses.

    robots.txt token; no separate HTTP request user-agent
    机器人行为
    Google-Extended is a robots.txt product-control token and does not affect Google Search inclusion or ranking.
    身份证据
    Not a request identity. Do not look for Google-Extended as a standalone HTTP user-agent in traffic logs.
    主要提供商来源 →
  • Applebot-Extended

    Apple

    control token

    Publisher control for whether Applebot-crawled content may be used to train Apple's general-purpose foundation models.

    robots.txt control token; does not crawl webpages
    机器人行为
    Applebot-Extended is configured in robots.txt but does not itself crawl webpages.
    身份证据
    Not a request identity. Traffic should be attributed to Applebot, not Applebot-Extended.
    主要提供商来源 →

开放网络语料库爬虫

1 entry
  • CCBot

    Common Crawl

    web corpus crawler

    Automated crawler that collects public web data for Common Crawl's open web-crawl repository.

    CCBot/2.0
    机器人行为
    Common Crawl documents robots.txt support for CCBot.
    身份证据
    Common Crawl publishes dedicated IP ranges and reverse-DNS guidance; the project also warns that clients can falsely claim the CCBot user-agent.
    主要提供商来源 →

不应将三个类别合并为一个“机器人”标签。

自动爬虫

Search and model-development crawlers operate automatically and are generally managed through provider-specific robots.txt policy.

用户触发的获取器

These requests happen because a user or agent asked a provider to retrieve or act on a page. Robots behavior can differ materially from automatic crawlers.

控制令牌

Google-Extended and Applebot-Extended are publisher controls, not standalone HTTP crawler identities. They should not be treated as traffic labels.

本目录确立了什么,又没有确立什么。

  • 它记录了在此页面被审查时,命名提供商所记录的身份和控制。
  • 它不确定携带这些字符串之一的请求是否合法;字符串可能被伪造。
  • 它不确定未识别的流量是否是人类。
  • 它不确定外部AI模型得出了什么结论,偏好,推荐,比较或决定。

提供商文档已检查 2026-08-29 · 在操作执行前重新检查链接的主要来源

Cartograph 正在围绕证据而非用户代理信任进行建设。

计划中的证据层旨在将声明的身份、验证证据、观察到的店面活动和未知项分开。Cartograph 不是机器人拦截器、WAF 或流量强制工具。