一份由来源支持的 AI 抓取器、代理和控制的地图。
此目录使用当前提供商文档将自动爬虫、用户触发的抓取工具和发布者控制令牌分开。它是解释候选 AI 代理流量和爬虫策略的参考——而不是声称每个列出的身份都访问过您的店面。
用户代理字符串并非身份证明,且提供商行为可能发生变化。在可能的情况下,使用提供商当前的 IP、DNS 或认证指南验证请求。未列出的身份不代表是人类流量。
已记录条目
16
已代表的运营商
7
已检查来源
Aug 29, 2026
用户触发的获取器和代理
5 entries- user triggered fetcher
ChatGPT-User
OpenAI
User-triggered fetcher used for certain actions in ChatGPT and Custom GPTs; it is not an automatic web crawler.
ChatGPT-User/1.0- 机器人行为
- OpenAI says robots.txt rules may not apply because these requests are user-initiated.
- 身份证据
- OpenAI publishes IP ranges for ChatGPT-User requests; the user-agent string remains spoofable.
- user triggered fetcher
Claude-User
Anthropic
Retrieves web content in response to user-initiated Claude requests.
Claude-User- 机器人行为
- Anthropic says its bots honor robots.txt directives, including Claude-User.
- 身份证据
- Anthropic's crawler guidance links a current source-IP list; do not trust the token alone.
- user triggered fetcher
Perplexity-User
Perplexity
User-triggered fetcher that may visit a page to help answer a Perplexity user's question; it is not used for web crawling or foundation-model training.
Perplexity-User/1.0- 机器人行为
- Perplexity says this fetcher generally ignores robots.txt because the request was initiated by a user.
- 身份证据
- Perplexity publishes IP ranges for Perplexity-User; combine the token and source IP rather than trusting the token alone.
- user triggered fetcher
Google-Agent
Google
Used by agents hosted on Google infrastructure to navigate the web and perform actions upon user request.
Google-Agent (browser-style user-agent string)- 机器人行为
- Google says user-triggered fetchers generally ignore robots.txt rules because the fetch was requested by a user.
- 身份证据
- Google publishes user-triggered-agent IP ranges and is experimenting with Web Bot Auth for the agent.bot.goog identity.
- user triggered fetcher
Amzn-User
Amazon
Supports user actions such as fetching current web information to answer Alexa queries on a user's behalf; Amazon says it is not used for generative-AI model training.
Amzn-User/0.1- 机器人行为
- Amazon says because Amzn-User actions can be user-initiated, it may not follow all robots.txt directives.
- 身份证据
- Amazon publishes IP addresses for Amzn-User; use source verification in addition to the token.
自动搜索爬虫
5 entries- search crawler
OAI-SearchBot
OpenAI
Automatic crawler used to surface websites in ChatGPT search results.
OAI-SearchBot/1.4 (documented example; version may change)- 机器人行为
- Managed with the OAI-SearchBot token in robots.txt.
- 身份证据
- Match the token with OpenAI's published SearchBot IP ranges; a user-agent string alone is not proof.
- search crawler
Claude-SearchBot
Anthropic
Navigates the web to improve the relevance and accuracy of Claude search responses.
Claude-SearchBot- 机器人行为
- Anthropic says its bots honor robots.txt directives.
- 身份证据
- Anthropic's crawler guidance links a current source-IP list; do not trust the token alone.
- search crawler
PerplexityBot
Perplexity
Automatic crawler designed to surface and link websites in Perplexity search results; Perplexity says it is not used to crawl content for foundation-model training.
PerplexityBot/1.0- 机器人行为
- Perplexity recommends managing PerplexityBot through robots.txt.
- 身份证据
- Perplexity publishes IP ranges for PerplexityBot; combine the token and source IP rather than trusting the token alone.
- search crawler
Applebot
Apple
Apple web crawler used for search experiences including Spotlight, Siri, and Safari; Apple says crawled data may also support foundation-model training and current-content context for AI outputs.
Applebot/<version> within Apple's documented browser-style user-agent format- 机器人行为
- Applebot respects standard robots.txt directives in general search crawls.
- 身份证据
- Apple documents reverse-DNS verification under *.applebot.apple.com and publishes Applebot IP CIDR ranges.
- search crawler
Amzn-SearchBot
Amazon
Amazon search crawler used to improve search experiences in Amazon products and services, including eligibility for experiences such as Alexa; Amazon says it is not used for generative-AI model training.
Amzn-SearchBot/0.1- 机器人行为
- Amazon documents robots.txt allow/disallow support for its crawlers.
- 身份证据
- Amazon publishes IP addresses for Amzn-SearchBot; use source verification in addition to the token.
模型开发爬虫
2 entries- model development crawler
GPTBot
OpenAI
Automatic crawler for content that may be used to train OpenAI generative AI foundation models.
GPTBot/1.4 (documented example; version may change)- 机器人行为
- Managed separately with the GPTBot token in robots.txt.
- 身份证据
- Match the token with OpenAI's published GPTBot IP ranges; a user-agent string alone is not proof.
- model development crawler
ClaudeBot
Anthropic
Collects public web content that could potentially contribute to Anthropic model training.
ClaudeBot- 机器人行为
- Anthropic says its bots honor robots.txt directives.
- 身份证据
- Anthropic's crawler guidance links a current source-IP list; do not trust the token alone.
通用提供商爬虫
1 entry- general crawler
Amazonbot
Amazon
Amazon crawler used to improve products and services; Amazon says collected content may also be used to train Amazon AI models.
Amazonbot/0.1- 机器人行为
- Amazon documents robots.txt allow/disallow support for its crawlers.
- 身份证据
- Amazon publishes IP addresses for Amazonbot; use source verification in addition to the token.
发布者控制令牌
2 entries- control token
Google-Extended
Google
Publisher control for whether Google-crawled content may be used for future Gemini model training and specified grounding uses.
robots.txt token; no separate HTTP request user-agent- 机器人行为
- Google-Extended is a robots.txt product-control token and does not affect Google Search inclusion or ranking.
- 身份证据
- Not a request identity. Do not look for Google-Extended as a standalone HTTP user-agent in traffic logs.
- control token
Applebot-Extended
Apple
Publisher control for whether Applebot-crawled content may be used to train Apple's general-purpose foundation models.
robots.txt control token; does not crawl webpages- 机器人行为
- Applebot-Extended is configured in robots.txt but does not itself crawl webpages.
- 身份证据
- Not a request identity. Traffic should be attributed to Applebot, not Applebot-Extended.
开放网络语料库爬虫
1 entry- web corpus crawler
CCBot
Common Crawl
Automated crawler that collects public web data for Common Crawl's open web-crawl repository.
CCBot/2.0- 机器人行为
- Common Crawl documents robots.txt support for CCBot.
- 身份证据
- Common Crawl publishes dedicated IP ranges and reverse-DNS guidance; the project also warns that clients can falsely claim the CCBot user-agent.
不应将三个类别合并为一个“机器人”标签。
自动爬虫
Search and model-development crawlers operate automatically and are generally managed through provider-specific robots.txt policy.
用户触发的获取器
These requests happen because a user or agent asked a provider to retrieve or act on a page. Robots behavior can differ materially from automatic crawlers.
控制令牌
Google-Extended and Applebot-Extended are publisher controls, not standalone HTTP crawler identities. They should not be treated as traffic labels.
本目录确立了什么,又没有确立什么。
- 它记录了在此页面被审查时,命名提供商所记录的身份和控制。
- 它不确定携带这些字符串之一的请求是否合法;字符串可能被伪造。
- 它不确定未识别的流量是否是人类。
- 它不确定外部AI模型得出了什么结论,偏好,推荐,比较或决定。
提供商文档已检查 2026-08-29 · 在操作执行前重新检查链接的主要来源
Cartograph 正在围绕证据而非用户代理信任进行建设。
计划中的证据层旨在将声明的身份、验证证据、观察到的店面活动和未知项分开。Cartograph 不是机器人拦截器、WAF 或流量强制工具。