Add my site

Crawler directory

Every crawler DataStated can name: the provider, the user-agent token, what the crawl is for, and where the provider publishes the addresses a request can be checked against. This page is generated from the registry the classifier runs on, so it is the same list your dashboard uses.

The registry names 50 agents from 17 providers and points at 18 published address lists. Each documentation link and each list was read from the provider's own site on September 7, 2026, and the token, purpose and list address were taken from that page. Where a provider publishes nothing, the entry says so, and hits from that crawler show as "could not be checked", never as confirmed. The address lists are re-read every six hours.

A token matches as a whole word, case-insensitively, so Googlebot-Image/1.0 is Googlebot-Image and never Googlebot, and NotGPTBot is nothing. A more specific token always sits ahead of the one it extends, and the first match wins. Anything automated that no entry names is kept as Other: well-known HTTP libraries and headless browsers under their own name, anything else under the bot word in its user agent.

Purposes

  • AI answers (ai_answer): a user asked an assistant something and it opened the page to answer.
  • Search indexing (search_index): building a search or answer index.
  • AI training (ai_training): collecting a corpus to train models.
  • Other (other): link previews, ads checks, tools, anything else.

The directory

Verification is published addresses when the provider publishes a list and a hit's address is checked against it, stated range when the provider gives its range in prose and that range is checked, reverse DNS when the provider offers only that (not looked up, so hits cannot be checked), and none when nothing is published.

ProviderAgentPurposeVerificationNote
OpenAIOAI-SearchBotSearch indexingPublished addresses: openai.com/searchbot.json
ChatGPT-UserAI answersPublished addresses: openai.com/chatgpt-user.json
GPTBotAI trainingPublished addresses: openai.com/gptbot.json
OAI-AdsBotOtherPublished addresses: openai.com/adsbot.jsonChecks pages submitted as ChatGPT ads.
AnthropicClaude-SearchBotSearch indexingPublished addresses: claude.com/crawling/bots.json
Claude-UserAI answersPublished addresses: claude.com/crawling/bots.json
ClaudeBotAI trainingPublished addresses: claude.com/crawling/bots.json
PerplexityPerplexity-UserAI answersPublished addresses: www.perplexity.ai/perplexity-user.json
PerplexityBotSearch indexingPublished addresses: www.perplexity.ai/perplexitybot.json
GoogleGooglebot-ImageSearch indexingPublished addresses: developers.google.com/static/crawling/ipranges/common-crawlers.json
Googlebot-VideoSearch indexingPublished addresses: developers.google.com/static/crawling/ipranges/common-crawlers.json
Googlebot-NewsSearch indexingPublished addresses: developers.google.com/static/crawling/ipranges/common-crawlers.json
GooglebotSearch indexingPublished addresses: developers.google.com/static/crawling/ipranges/common-crawlers.json
Google-ExtendedAI trainingPublished addresses: developers.google.com/static/crawling/ipranges/common-crawlers.jsonA robots.txt control token for Gemini training; Google says it does not appear as its own user agent, so hits under this name are unexpected.
GoogleOther-ImageOtherPublished addresses: developers.google.com/static/crawling/ipranges/common-crawlers.json
GoogleOther-VideoOtherPublished addresses: developers.google.com/static/crawling/ipranges/common-crawlers.json
GoogleOtherOtherPublished addresses: developers.google.com/static/crawling/ipranges/common-crawlers.jsonGoogle's generic crawler for product teams; purpose unstated per fetch.
Google-CloudVertexBotOtherPublished addresses: developers.google.com/static/crawling/ipranges/common-crawlers.jsonCrawls a site at its owner's request to build Vertex AI agents.
Storebot-GoogleOtherPublished addresses: developers.google.com/static/crawling/ipranges/common-crawlers.json
Google-InspectionToolOtherPublished addresses: developers.google.com/static/crawling/ipranges/common-crawlers.json
Google-GeminiNotebookAI answersPublished addresses: developers.google.com/static/crawling/ipranges/user-triggered-fetchers.json
Google-AgentAI answersPublished addresses: developers.google.com/static/crawling/ipranges/user-triggered-agents.jsonAgents on Google infrastructure acting on a user's request.
FeedFetcher-GoogleOtherPublished addresses: developers.google.com/static/crawling/ipranges/user-triggered-fetchers.json
Google-Read-AloudOtherPublished addresses: developers.google.com/static/crawling/ipranges/user-triggered-fetchers.json
BingbingbotSearch indexingPublished addresses: www.bing.com/toolbox/bingbot.json
BingPreviewOtherReverse DNS under search.msn.com, not looked upBing's crawler list page is script-rendered and could not be read on 2026-09-07; the token is the one Bing prints in its user agents. Bing says its crawlers reverse-resolve to search.msn.com.
AdIdxBotOtherReverse DNS under search.msn.com, not looked upBing's crawler list page is script-rendered and could not be read on 2026-09-07; the token is the one Bing prints in its user agents. Bing says its crawlers reverse-resolve to search.msn.com.
AppleApplebot-ExtendedAI trainingNone publishedA robots.txt control token; Apple says it does not crawl on its own, so hits under this name are unexpected.
ApplebotSearch indexingPublished addresses: search.developer.apple.com/applebot.jsonApple also says Applebot reverse-resolves to *.applebot.apple.com.
Metameta-externalfetcherAI answersNone publishedMeta's crawler page names the agent and its purpose but publishes no address list.
meta-externalagentAI trainingNone publishedMeta's crawler page names the agent and its purpose but publishes no address list.
meta-webindexerSearch indexingNone publishedMeta's crawler page names the agent and its purpose but publishes no address list.
meta-externaladsOtherNone publishedMeta's crawler page names the agent and its purpose but publishes no address list.
facebookexternalhitOtherNone publishedLink previews for content shared on Facebook, Instagram and Messenger. Meta's crawler page names the agent and its purpose but publishes no address list.
WhatsAppOtherNone publishedLink previews; the agent is WhatsApp/<version> followed by A, I or N.
AmazonAmazonbotAI trainingPublished addresses: developer.amazon.com/amazonbot/ip-addresses/Amazon says the crawl may train its AI models. The address list is published inside a web page, not as a file; the worker scans it out of the page.
Common CrawlCCBotAI trainingPublished addresses: index.commoncrawl.org/ccbot.jsonAn open crawl archive that model builders train from.
ByteDanceBytespiderAI trainingNone publishedNo reachable official documentation on 2026-09-07: the agent's own link (zhanzhang.toutiao.com) does not resolve outside China.
DuckDuckGoDuckAssistBotAI answersPublished addresses: duckduckgo.com/duckassistbot.json
DuckDuckBotSearch indexingPublished addresses: duckduckgo.com/duckduckbot.json
MistralMistralAI-UserAI answersPublished addresses: mistral.ai/mistralai-user-ips.json
MistralAI-IndexSearch indexingPublished addresses: mistral.ai/mistralai-index-ips.json
MistralAI-TrainingAI trainingNone publishedMistral publishes address lists for its other two agents but not this one.
You.comYouBotSearch indexingStated range 68.67.112.0/24You.com states the range in prose and also signs requests (Web Bot Auth); only the range is checked here.
Coherecohere-training-data-crawlerAI trainingNone publishedCohere's own page lists no active crawler ("N/A") and says it does not crawl to train models at this time; this token is the one seen in server logs.
cohere-aiAI trainingNone publishedCohere's own page lists no active crawler ("N/A") and says it does not crawl to train models at this time; this token is the one seen in server logs.
DiffbotDiffbot-UserOtherNone publishedA person browsing a URL through Diffbot software.
DiffbotSearch indexingNone publishedDiffbot says the crawl builds its Knowledge Graph and search; no address list is published.
YahooSlurpSearch indexingNone publishedYahoo's page names the agent only.
TelegramTelegramBotOtherNone publishedLink previews. Telegram publishes no page describing the agent; the string is TelegramBot (like TwitterBot).

Questions? Email us at hello@datastated.com.