Crawler directory
Every crawler DataStated can name: the provider, the user-agent token, what the crawl is for, and where the provider publishes the addresses a request can be checked against. This page is generated from the registry the classifier runs on, so it is the same list your dashboard uses.
The registry names 50 agents from 17 providers and points at 18 published address lists. Each documentation link and each list was read from the provider's own site on September 7, 2026, and the token, purpose and list address were taken from that page. Where a provider publishes nothing, the entry says so, and hits from that crawler show as "could not be checked", never as confirmed. The address lists are re-read every six hours.
A token matches as a whole word, case-insensitively, so Googlebot-Image/1.0 is Googlebot-Image and never Googlebot, and NotGPTBot is nothing. A more specific token always sits ahead of the one it extends, and the first match wins. Anything automated that no entry names is kept as Other: well-known HTTP libraries and headless browsers under their own name, anything else under the bot word in its user agent.
Purposes
- AI answers (
ai_answer): a user asked an assistant something and it opened the page to answer. - Search indexing (
search_index): building a search or answer index. - AI training (
ai_training): collecting a corpus to train models. - Other (
other): link previews, ads checks, tools, anything else.
The directory
Verification is published addresses when the provider publishes a list and a hit's address is checked against it, stated range when the provider gives its range in prose and that range is checked, reverse DNS when the provider offers only that (not looked up, so hits cannot be checked), and none when nothing is published.
| Provider | Agent | Purpose | Verification | Note |
|---|---|---|---|---|
| OpenAI | OAI-SearchBot | Search indexing | Published addresses: openai.com/searchbot.json | |
ChatGPT-User | AI answers | Published addresses: openai.com/chatgpt-user.json | ||
GPTBot | AI training | Published addresses: openai.com/gptbot.json | ||
OAI-AdsBot | Other | Published addresses: openai.com/adsbot.json | Checks pages submitted as ChatGPT ads. | |
| Anthropic | Claude-SearchBot | Search indexing | Published addresses: claude.com/crawling/bots.json | |
Claude-User | AI answers | Published addresses: claude.com/crawling/bots.json | ||
ClaudeBot | AI training | Published addresses: claude.com/crawling/bots.json | ||
| Perplexity | Perplexity-User | AI answers | Published addresses: www.perplexity.ai/perplexity-user.json | |
PerplexityBot | Search indexing | Published addresses: www.perplexity.ai/perplexitybot.json | ||
Googlebot-Image | Search indexing | Published addresses: developers.google.com/static/crawling/ipranges/common-crawlers.json | ||
Googlebot-Video | Search indexing | Published addresses: developers.google.com/static/crawling/ipranges/common-crawlers.json | ||
Googlebot-News | Search indexing | Published addresses: developers.google.com/static/crawling/ipranges/common-crawlers.json | ||
Googlebot | Search indexing | Published addresses: developers.google.com/static/crawling/ipranges/common-crawlers.json | ||
Google-Extended | AI training | Published addresses: developers.google.com/static/crawling/ipranges/common-crawlers.json | A robots.txt control token for Gemini training; Google says it does not appear as its own user agent, so hits under this name are unexpected. | |
GoogleOther-Image | Other | Published addresses: developers.google.com/static/crawling/ipranges/common-crawlers.json | ||
GoogleOther-Video | Other | Published addresses: developers.google.com/static/crawling/ipranges/common-crawlers.json | ||
GoogleOther | Other | Published addresses: developers.google.com/static/crawling/ipranges/common-crawlers.json | Google's generic crawler for product teams; purpose unstated per fetch. | |
Google-CloudVertexBot | Other | Published addresses: developers.google.com/static/crawling/ipranges/common-crawlers.json | Crawls a site at its owner's request to build Vertex AI agents. | |
Storebot-Google | Other | Published addresses: developers.google.com/static/crawling/ipranges/common-crawlers.json | ||
Google-InspectionTool | Other | Published addresses: developers.google.com/static/crawling/ipranges/common-crawlers.json | ||
Google-GeminiNotebook | AI answers | Published addresses: developers.google.com/static/crawling/ipranges/user-triggered-fetchers.json | ||
Google-Agent | AI answers | Published addresses: developers.google.com/static/crawling/ipranges/user-triggered-agents.json | Agents on Google infrastructure acting on a user's request. | |
FeedFetcher-Google | Other | Published addresses: developers.google.com/static/crawling/ipranges/user-triggered-fetchers.json | ||
Google-Read-Aloud | Other | Published addresses: developers.google.com/static/crawling/ipranges/user-triggered-fetchers.json | ||
| Bing | bingbot | Search indexing | Published addresses: www.bing.com/toolbox/bingbot.json | |
BingPreview | Other | Reverse DNS under search.msn.com, not looked up | Bing's crawler list page is script-rendered and could not be read on 2026-09-07; the token is the one Bing prints in its user agents. Bing says its crawlers reverse-resolve to search.msn.com. | |
AdIdxBot | Other | Reverse DNS under search.msn.com, not looked up | Bing's crawler list page is script-rendered and could not be read on 2026-09-07; the token is the one Bing prints in its user agents. Bing says its crawlers reverse-resolve to search.msn.com. | |
| Apple | Applebot-Extended | AI training | None published | A robots.txt control token; Apple says it does not crawl on its own, so hits under this name are unexpected. |
Applebot | Search indexing | Published addresses: search.developer.apple.com/applebot.json | Apple also says Applebot reverse-resolves to *.applebot.apple.com. | |
| Meta | meta-externalfetcher | AI answers | None published | Meta's crawler page names the agent and its purpose but publishes no address list. |
meta-externalagent | AI training | None published | Meta's crawler page names the agent and its purpose but publishes no address list. | |
meta-webindexer | Search indexing | None published | Meta's crawler page names the agent and its purpose but publishes no address list. | |
meta-externalads | Other | None published | Meta's crawler page names the agent and its purpose but publishes no address list. | |
facebookexternalhit | Other | None published | Link previews for content shared on Facebook, Instagram and Messenger. Meta's crawler page names the agent and its purpose but publishes no address list. | |
WhatsApp | Other | None published | Link previews; the agent is WhatsApp/<version> followed by A, I or N. | |
| Amazon | Amazonbot | AI training | Published addresses: developer.amazon.com/amazonbot/ip-addresses/ | Amazon says the crawl may train its AI models. The address list is published inside a web page, not as a file; the worker scans it out of the page. |
| Common Crawl | CCBot | AI training | Published addresses: index.commoncrawl.org/ccbot.json | An open crawl archive that model builders train from. |
| ByteDance | Bytespider | AI training | None published | No reachable official documentation on 2026-09-07: the agent's own link (zhanzhang.toutiao.com) does not resolve outside China. |
| DuckDuckGo | DuckAssistBot | AI answers | Published addresses: duckduckgo.com/duckassistbot.json | |
DuckDuckBot | Search indexing | Published addresses: duckduckgo.com/duckduckbot.json | ||
| Mistral | MistralAI-User | AI answers | Published addresses: mistral.ai/mistralai-user-ips.json | |
MistralAI-Index | Search indexing | Published addresses: mistral.ai/mistralai-index-ips.json | ||
MistralAI-Training | AI training | None published | Mistral publishes address lists for its other two agents but not this one. | |
| You.com | YouBot | Search indexing | Stated range 68.67.112.0/24 | You.com states the range in prose and also signs requests (Web Bot Auth); only the range is checked here. |
| Cohere | cohere-training-data-crawler | AI training | None published | Cohere's own page lists no active crawler ("N/A") and says it does not crawl to train models at this time; this token is the one seen in server logs. |
cohere-ai | AI training | None published | Cohere's own page lists no active crawler ("N/A") and says it does not crawl to train models at this time; this token is the one seen in server logs. | |
| Diffbot | Diffbot-User | Other | None published | A person browsing a URL through Diffbot software. |
Diffbot | Search indexing | None published | Diffbot says the crawl builds its Knowledge Graph and search; no address list is published. | |
| Yahoo | Slurp | Search indexing | None published | Yahoo's page names the agent only. |
| Telegram | TelegramBot | Other | None published | Link previews. Telegram publishes no page describing the agent; the string is TelegramBot (like TwitterBot). |
Questions? Email us at hello@datastated.com.