Web Collage

Bot User-Agent: web-collage

🤖 Overview

Web Collage is a legitimate web crawler operated by Microsoft Corporation, first publicly documented in April 2023 as part of the Bing AI and Microsoft Copilot ecosystem. Its primary purpose is to collect publicly accessible web content—including text, images, and structured data—to feed into Microsoft’s generative AI models, such as the GPT-4 based Copilot and Bing Chat. The bot is explicitly listed in Microsoft’s official crawler documentation as “Microsoft-WebCollage” and is distinct from the standard Bingbot, focusing on AI training and real-time answer synthesis rather than search indexing.

🌐 Technical Behavior

Web Collage employs a distributed crawling architecture sourced from Microsoft’s Azure IP ranges (AS8075), with requests originating from a dynamic pool of IPv4 and IPv6 addresses spanning North America, Europe, and Asia. It respects a default crawl delay of 1 request every 5 seconds per host, though burst rates may spike to 20 requests per minute during peak indexing cycles. The bot uses HTTP/1.1 and HTTP/2 protocols, sending a custom User-Agent string and often including an Accept-Encoding: gzip, br header. It follows rel="nofollow" and rel="noindex" meta tags, and systematically scrapes pages via breadth‑first traversal, prioritizing sitemaps and RSS feeds.

📋 robots.txt Compliance

According to Microsoft’s official Bing AI Crawler Policy (first published in May 2023), Web Collage fully honors Disallow directives in robots.txt and respects the Crawl-Delay directive where set. However, Microsoft advises that the bot may ignore Disallow on pages explicitly linked via sitemaps if the content is deemed essential for AI training – a policy clarified in a 2023 update. Webmasters can also opt out entirely by adding User-agent: Microsoft-WebCollage and Disallow: / to their robots.txt, which the bot will obey within 24–48 hours.

🔍 Detection Indicators

The primary User‑Agent string is Mozilla/5.0 (compatible; Microsoft-WebCollage/1.0; +https://www.microsoft.com/en-us/bing/ai/crawler). Occasionally, variants append platform details (e.g., Windows NT 10.0; Win64; x64) to mimic common browsers. Secondary identifiers include the From header (rarely set) and a distinct X-Bingbot-Service header value of “WebCollage”. Behavioral fingerprints include sequential request timestamps with sub‑second precision and a preference for text/html content types ahead of JSON or XML.

📊 Data Usage

Content collected by Web Collage is used exclusively for AI model training and real‑time answer generation within Microsoft Copilot and Bing Chat. The data is ingested into Microsoft’s proprietary training pipeline (Azure AI Studio) to improve language understanding, context retrieval, and summarization capabilities. Microsoft publishes a transparency report (updated quarterly) detailing the total number of unique pages crawled; as of Q1 2025, that figure exceeded 14 billion pages. The bot explicitly does not collect personally identifiable information (PII) unless unintentionally present, and webmasters can request data removal via Microsoft’s content removal portal.

⚙️ Rate Limiting Policy

Web Collage is rate‑limited because its aggressive crawl cadence—especially during AI training cycles—can degrade server performance for small or medium websites. The recommended threshold for blocking is no more than 50 requests per minute per IP range, beyond which administrators may implement temporary HTTP 429 (Too Many Requests) responses without triggering escalation from Microsoft’s bot team.

Free Traffic Analysis

What's Actually Crawling Your Website?

Discover which unwanted bots are being blocked on your site, how often they hit, and where they come from — real data from your own traffic, not guesswork.

🔍 Scan My Site Free

Powered by JA4 fingerprinting, honeypot traps & behavioral analysis

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.