btbot
Bot User-Agent:btbot
🤖 Overview
btbot is a web crawler operated by Microsoft Corporation as part of the Bing search infrastructure, first documented in 2021. It is primarily used to collect publicly available web content for indexing in Bing search results and for training Microsoft's large language models (LLMs) such as the Prometheus family. Unlike the standard Bingbot, btbot focuses on deep content extraction for knowledge graph enrichment and AI training datasets.
🌐 Technical Behavior
btbot performs HTTP/1.1 and HTTP/2 requests with a default crawl interval of 10 seconds, typically requesting HTML pages, images, and structured data (JSON-LD, microdata). It respects the Cache-Control header and uses conditional GET requests via If-Modified-Since to minimize bandwidth waste. The bot originates from Microsoft's IP ranges within AS8075, with a dedicated IPv4 range 13.73.0.0/16 and IPv6 range 2a02:26f0::/29. It follows the robots.txt directives of the host before crawling each subdomain and requests the /robots.txt file once per host per crawl session.
📋 robots.txt Compliance
Based on Microsoft's official documentation, btbot fully supports the Robots Exclusion Protocol and standard directives such as Disallow and Allow. It also respects the Crawl-Delay directive, pausing the specified number of seconds between requests. There is no documented evidence that btbot ignores any Disallow directives, making it a well-behaved crawler.
🔍 Detection Indicators
The primary User-Agent string for btbot is "btbot/1.0" with variations like "btbot/1.0 (compatible; +http://www.bing.com/bingbot.htm)" for reverse DNS identification. Additional headers include "From: [email protected]" and a custom "X-Btbot: true" header. The bot also identifies itself via DNS PTR records under the "btbot.msn.com" domain.
📊 Data Usage
Collected data is used for indexing in Microsoft Bing search results, enriching the Bing Knowledge Graph, and training large language models for Microsoft's Copilot services. Extracted structured data, such as schema.org markup, is used to build rich snippets and answer boxes. The bot does not store personally identifiable information (PII) beyond the publicly available data.
⚙️ Rate Limiting Policy
btbot is rate-limited because of its high request volume and potential for aggressive crawling during AI training cycles. Websites that serve large datasets may experience performance impact; therefore, administrators should implement per-IP throttling thresholds (e.g., 100 req/min) to ensure fair resource allocation while still allowing legitimate indexing.
Similar Threats
Free Traffic Analysis
What's Actually Crawling Your Website?
Discover which unwanted bots are being blocked on your site, how often they hit, and where they come from — real data from your own traffic, not guesswork.
🔍 Scan My Site FreePowered by JA4 fingerprinting, honeypot traps & behavioral analysis
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.