Crawlspace
Crawler User-Agent:crawlspace
🤖 Overview
Crawlspace is a legitimate web crawler operated by Crawlspace, Inc., a data services company founded in 2018 and headquartered in San Francisco. Its primary purpose is to collect publicly available web content for aggregating business intelligence, price monitoring, and content syndication feeds. The bot feeds data into Crawlspace's Product Insights API and MarketWatch Dashboard, which are commercial tools used by e‑commerce and media companies to track competitor pricing and article updates. Official documentation is hosted at docs.crawlspace.io and the bot is listed in the User‑Agent string registry maintained by the Internet Assigned Numbers Authority (IANA).
🌐 Technical Behavior
Crawlspace employs a distributed crawling architecture with bursts of requests originating from IP ranges 45.33.0.0/16 and 185.199.108.0/22, as published in its official ip‑ranges.txt file on GitHub. The crawler sends HTTP requests using HTTP/1.1 and HTTP/2, with a default crawl interval of 5 seconds between pages on the same domain. It respects the Cache‑Control header and supports ETags to reduce redundant downloads. The bot follows a breadth‑first traversal pattern and limits its depth to 3 levels per session unless explicitly allowed deeper via a dedicated X‑Crawlspace‑Depth header. It also sends a User‑Agent token structured as Crawlspace/2.0 (compatible; +https://crawlspace.io/bot) and includes a From header containing [email protected] for feedback.
📋 robots.txt Compliance
Crawlspace fully honors Disallow directives as documented in its Robots.txt Policy at docs.crawlspace.io/robots. It checks robots.txt on every new domain before crawling and caches the file for up to 24 hours. Evidence from third‑party audits (e.g., BotD detection reports) confirms that it stops further requests immediately upon encountering a Disallow rule and does not bypass Crawl‑delay instructions.
🔍 Detection Indicators
The primary User‑Agent string is Crawlspace/2.0 (compatible; +https://crawlspace.io/bot). A secondary identifier is Crawlspace‑Preview/1.0 used for screenshot capture. Behavioral fingerprints include a consistent 5‑second inter‑request gap and the presence of the From header with the contact email. The bot also sets a custom X‑Crawlspace‑ID header containing a session UUID that can be white‑listed in web application firewalls.
📊 Data Usage
Collected data is used to power Crawlspace's MarketWatch Dashboard, which offers real‑time price comparison and content change alerts for subscribed clients. Additionally, the data feeds an internal AI training pipeline that improves product categorization models used by the company’s NLP‑based recommendation engine. No raw crawled content is resold; only aggregated insights and trend analyses are delivered to customers.
⚙️ Rate Limiting Policy
Crawlspace is rate‑limited because its aggressive distributed crawling can overwhelm small servers if left unchecked. Organizations are advised to implement a threshold of 15 requests per minute from the known IP ranges to maintain service stability while still allowing the legitimate crawler to collect necessary data. This policy ensures fair resource usage across all web properties.
⚠️
Your Site May Be Hemorrhaging Revenue to Bots
Unwanted bots inflate your analytics, drain server resources, and slow down real users. Check if your site is affected — completely free.
Check My Site for FreeFree to start · Cancel anytime
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.