collapsarweb

Bot User-Agent: collapsarweb

🤖 Overview

Collapsarweb is a web crawler operated by Collapsar Technology Co., Ltd., a Chinese data analytics and AI company headquartered in Beijing. First documented in 2022, its primary purpose is to collect publicly accessible web content for training proprietary machine learning models and powering the company’s search‑indexing service, CollapsarSearch (a vertical search engine for technology and academic domains). The bot is explicitly listed in the user‑agent database of major web servers and has been observed in server logs since early 2023.

🌐 Technical Behavior

Collapsarweb follows a breadth‑first crawl strategy, typically issuing requests at a rate of 1–3 per second per source IP, with bursts of up to 10 requests per second during initial discovery phases. It uses HTTP/1.1 and HTTP/2 protocols, including support for gzip and deflate compression. Officially, its IP ranges belong to the ASN AS45090 (China Telecom) and AS37963 (China Unicom), with addresses primarily located in Beijing and Shanghai. The crawler respects If‑Modified‑Since headers to avoid re‑downloading unchanged content and includes a Referer header pointing to the parent page it is crawling. It does not execute JavaScript and will not follow rel="nofollow" links.

📋 robots.txt Compliance

According to Collapsar Technology’s official documentation (published at https://www.collapsar.com/bot/robots), Collapsarweb fully honors Disallow directives in robots.txt and also respects Crawl‑Delay instructions. Third‑party server log audits (e.g., the 2023 report by BotDrill Research) confirm that the bot does not access paths disallowed in robots.txt and pauses between requests when a delay is specified.

🔍 Detection Indicators

The primary User‑Agent string is: Mozilla/5.0 (compatible; Collapsarweb/1.0; +http://www.collapsar.com/bot.asp). A secondary string CollapsarBot/2.0 is used for mobile‑optimised crawls. The bot also sends a custom HTTP header X‑Collapsar‑Crawl: v1 and a From header containing a contact email ([email protected]). The IP’s reverse DNS resolves to *.crawl.collapsar.com.

📊 Data Usage

Collected data is used for training the company’s Collapsar‑NLP suite of language models (documented at https://github.com/CollapsarTech/collapsar-nlp), for populating the CollapsarSearch index, and for internal analytics on web technology adoption trends. The data is stored on encrypted servers in Beijing and is never sold to third parties; it is used exclusively for product improvement.

⚙️ Rate Limiting Policy

Because Collapsarweb can generate sustained traffic during large‑scale index refreshes (often exceeding 50,000 requests per day per site), it is rate‑limited to prevent resource exhaustion. Most webmasters apply a threshold of 5 requests per second per IP; the bot’s own documentation advises respecting a Crawl‑Delay: 2 directive to maintain fair use.

Free Bot Analysis

Is Your Site Under Bot Attack Right Now?

Find out exactly how much of your traffic is automated — and which bots are draining your bandwidth and skewing your analytics.

Run Free Bot Scan →

No credit card required  ·  Results in minutes

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.