הפניה מתועדת על ידי ספק

מפה מגובה מקור של סורקי AI, סוכנים ובקרות.

מדריך זה מפריד בין סורקים אוטומטיים, מביאים המופעלים על ידי משתמשים, ואסימוני בקרה של מפרסמים באמצעות תיעוד ספק נוכחי. הוא משמש כנקודת התייחסות לפרשנות תעבורת סוכני AI וטקטיקות זחילה — ולא כטענה שכל זהות המופיעה בו ביקרה בחזית החנות שלך.

מחרוזת User-agent אינה הוכחה לזהות, והתנהגות הספק יכולה להשתנות. יש לוודא בקשות באמצעות ה-IP הנוכחי של הספק, DNS או הנחיות אימות היכן שזמינות. היעדר זהות רשומה אינו מוכיח תעבורה אנושית.

ערכים מתועדים

16

מפעילי מערכת מיוצגים

7

מקורות שנבדקו

Aug 29, 2026

מביאים וסוכנים המופעלים על ידי המשתמש

5 כניסהies
  • ChatGPT-User

    OpenAI

    user triggered fetcher

    User-triggered fetcher used for certain actions in ChatGPT and Custom GPTs; it is not an automatic web crawler.

    ChatGPT-User/1.0
    התנהגות רובוטים
    OpenAI says robots.txt rules may not apply because these requests are user-initiated.
    ראיות זהות
    OpenAI publishes IP ranges for ChatGPT-User requests; the user-agent string remains spoofable.
    מקור ספק ראשי ←
  • Claude-User

    Anthropic

    user triggered fetcher

    Retrieves web content in response to user-initiated Claude requests.

    Claude-User
    התנהגות רובוטים
    Anthropic says its bots honor robots.txt directives, including Claude-User.
    ראיות זהות
    Anthropic's crawler guidance links a current source-IP list; do not trust the token alone.
    מקור ספק ראשי ←
  • Perplexity-User

    Perplexity

    user triggered fetcher

    User-triggered fetcher that may visit a page to help answer a Perplexity user's question; it is not used for web crawling or foundation-model training.

    Perplexity-User/1.0
    התנהגות רובוטים
    Perplexity says this fetcher generally ignores robots.txt because the request was initiated by a user.
    ראיות זהות
    Perplexity publishes IP ranges for Perplexity-User; combine the token and source IP rather than trusting the token alone.
    מקור ספק ראשי ←
  • Google-Agent

    Google

    user triggered fetcher

    Used by agents hosted on Google infrastructure to navigate the web and perform actions upon user request.

    Google-Agent (browser-style user-agent string)
    התנהגות רובוטים
    Google says user-triggered fetchers generally ignore robots.txt rules because the fetch was requested by a user.
    ראיות זהות
    Google publishes user-triggered-agent IP ranges and is experimenting with Web Bot Auth for the agent.bot.goog identity.
    מקור ספק ראשי ←
  • Amzn-User

    Amazon

    user triggered fetcher

    Supports user actions such as fetching current web information to answer Alexa queries on a user's behalf; Amazon says it is not used for generative-AI model training.

    Amzn-User/0.1
    התנהגות רובוטים
    Amazon says because Amzn-User actions can be user-initiated, it may not follow all robots.txt directives.
    ראיות זהות
    Amazon publishes IP addresses for Amzn-User; use source verification in addition to the token.
    מקור ספק ראשי ←

סורקי חיפוש אוטומטיים

5 כניסהies
  • OAI-SearchBot

    OpenAI

    search crawler

    Automatic crawler used to surface websites in ChatGPT search results.

    OAI-SearchBot/1.4 (documented example; version may change)
    התנהגות רובוטים
    Managed with the OAI-SearchBot token in robots.txt.
    ראיות זהות
    Match the token with OpenAI's published SearchBot IP ranges; a user-agent string alone is not proof.
    מקור ספק ראשי ←
  • Claude-SearchBot

    Anthropic

    search crawler

    Navigates the web to improve the relevance and accuracy of Claude search responses.

    Claude-SearchBot
    התנהגות רובוטים
    Anthropic says its bots honor robots.txt directives.
    ראיות זהות
    Anthropic's crawler guidance links a current source-IP list; do not trust the token alone.
    מקור ספק ראשי ←
  • PerplexityBot

    Perplexity

    search crawler

    Automatic crawler designed to surface and link websites in Perplexity search results; Perplexity says it is not used to crawl content for foundation-model training.

    PerplexityBot/1.0
    התנהגות רובוטים
    Perplexity recommends managing PerplexityBot through robots.txt.
    ראיות זהות
    Perplexity publishes IP ranges for PerplexityBot; combine the token and source IP rather than trusting the token alone.
    מקור ספק ראשי ←
  • Applebot

    Apple

    search crawler

    Apple web crawler used for search experiences including Spotlight, Siri, and Safari; Apple says crawled data may also support foundation-model training and current-content context for AI outputs.

    Applebot/<version> within Apple's documented browser-style user-agent format
    התנהגות רובוטים
    Applebot respects standard robots.txt directives in general search crawls.
    ראיות זהות
    Apple documents reverse-DNS verification under *.applebot.apple.com and publishes Applebot IP CIDR ranges.
    מקור ספק ראשי ←
  • Amzn-SearchBot

    Amazon

    search crawler

    Amazon search crawler used to improve search experiences in Amazon products and services, including eligibility for experiences such as Alexa; Amazon says it is not used for generative-AI model training.

    Amzn-SearchBot/0.1
    התנהגות רובוטים
    Amazon documents robots.txt allow/disallow support for its crawlers.
    ראיות זהות
    Amazon publishes IP addresses for Amzn-SearchBot; use source verification in addition to the token.
    מקור ספק ראשי ←

סורקי פיתוח מודלים

2 כניסהies
  • GPTBot

    OpenAI

    model development crawler

    Automatic crawler for content that may be used to train OpenAI generative AI foundation models.

    GPTBot/1.4 (documented example; version may change)
    התנהגות רובוטים
    Managed separately with the GPTBot token in robots.txt.
    ראיות זהות
    Match the token with OpenAI's published GPTBot IP ranges; a user-agent string alone is not proof.
    מקור ספק ראשי ←
  • ClaudeBot

    Anthropic

    model development crawler

    Collects public web content that could potentially contribute to Anthropic model training.

    ClaudeBot
    התנהגות רובוטים
    Anthropic says its bots honor robots.txt directives.
    ראיות זהות
    Anthropic's crawler guidance links a current source-IP list; do not trust the token alone.
    מקור ספק ראשי ←

סורקים כלליים של ספקים

1 כניסהy
  • Amazonbot

    Amazon

    general crawler

    Amazon crawler used to improve products and services; Amazon says collected content may also be used to train Amazon AI models.

    Amazonbot/0.1
    התנהגות רובוטים
    Amazon documents robots.txt allow/disallow support for its crawlers.
    ראיות זהות
    Amazon publishes IP addresses for Amazonbot; use source verification in addition to the token.
    מקור ספק ראשי ←

אסימוני בקרת מפרסם

2 כניסהies
  • Google-Extended

    Google

    control token

    Publisher control for whether Google-crawled content may be used for future Gemini model training and specified grounding uses.

    robots.txt token; no separate HTTP request user-agent
    התנהגות רובוטים
    Google-Extended is a robots.txt product-control token and does not affect Google Search inclusion or ranking.
    ראיות זהות
    Not a request identity. Do not look for Google-Extended as a standalone HTTP user-agent in traffic logs.
    מקור ספק ראשי ←
  • Applebot-Extended

    Apple

    control token

    Publisher control for whether Applebot-crawled content may be used to train Apple's general-purpose foundation models.

    robots.txt control token; does not crawl webpages
    התנהגות רובוטים
    Applebot-Extended is configured in robots.txt but does not itself crawl webpages.
    ראיות זהות
    Not a request identity. Traffic should be attributed to Applebot, not Applebot-Extended.
    מקור ספק ראשי ←

סורקי קורפוס אינטרנט פתוח

1 כניסהy
  • CCBot

    Common Crawl

    web corpus crawler

    Automated crawler that collects public web data for Common Crawl's open web-crawl repository.

    CCBot/2.0
    התנהגות רובוטים
    Common Crawl documents robots.txt support for CCBot.
    ראיות זהות
    Common Crawl publishes dedicated IP ranges and reverse-DNS guidance; the project also warns that clients can falsely claim the CCBot user-agent.
    מקור ספק ראשי ←

אין לאחד שלוש קטגוריות שונות לתווית אחת של "בוט".

זחלנים אוטומטיים

Search and model-development crawlers operate automatically and are generally managed through provider-specific robots.txt policy.

מאחזרי מידע שהופעלו על ידי המשתמש

These requests happen because a user or agent asked a provider to retrieve or act on a page. Robots behavior can differ materially from automatic crawlers.

אסימוני בקרה

Google-Extended and Applebot-Extended are publisher controls, not standalone HTTP crawler identities. They should not be treated as traffic labels.

מה ספרייה זו קובעת – ומה לא.

  • הוא מתעד זהויות ובקרות שתועדו על ידי הספקים הנקובים בעת סקירת עמוד זה.
  • זה לא קובע שבקשה הנושאת אחת מהמחרוזות הללו היא לגיטימית; ניתן לזייף מחרוזות.
  • זה לא קובע שתעבורה לא מזוהה היא אנושית.
  • זה לא קובע מה מודל AI חיצוני הסיק, העדיף, המליץ, השווה או החליט.

תיעוד ספק נבדק 2026-08-29 · בדוק מחדש את המקור העיקרי המקושר לפני אכיפה תפעולית

Cartograph נבנה סביב ראיות, לא סביב אמון בסוכן המשתמש.

שכבת הראיות המתוכננת נועדה להפריד זהות מוצהרת, ראיות אימות, פעילות חזית חנות שנצפתה וגורמים לא ידועים. Cartograph אינה חוסם בוטים, WAF או כלי לאכיפת תעבורה.