sqworm
Bot User-Agent:sqworm
🤖 Overview
Sqworm is a legitimate web crawler operated by Sqworm Inc., a data analytics company headquartered in San Francisco, California. According to public documentation on Sqworm’s official website (sqworm.com/crawler), the bot is designed to systematically index publicly accessible web content for use in AI training datasets and natural language processing research. It collects text, metadata, and structural data from web pages, which are then aggregated into large-scale corpora for improving machine learning models, particularly for language understanding and generation tasks.
🌐 Technical Behavior
Sqworm follows a distributed crawling architecture using multiple concurrent threads, with requests originating from IP ranges registered under ASN 396356 (Sqworm Inc., data center in Ashburn, VA). The bot sends HTTP requests at a rate of approximately 50–100 requests per second per crawling instance, but it dynamically reduces its speed when it detects server load spikes. It supports both HTTP/1.1 and HTTP/2 protocols, and it fetches pages with a default timeout of 30 seconds. Sqworm respects Cache-Control and ETag headers to avoid redundant downloads, and it sends a custom From header with the contact email [email protected]. The crawler does not execute JavaScript, CSS, or other client-side resources; it only parses raw HTML and linked text content.
📋 robots.txt Compliance
Sqworm explicitly honors robots.txt Disallow directives. The official documentation states that the bot checks robots.txt before each crawl and caches the file for up to 24 hours. It also respects Crawl-Delay directives when present, adhering to the specified delay in seconds. Tests by third-party researchers (such as those documented on bot-tests.io) have confirmed that Sqworm consistently stops crawling paths listed in Disallow rules and does not ignore them under any known circumstance.
🔍 Detection Indicators
The primary User-Agent string for Sqworm is Mozilla/5.0 (compatible; Sqworm/1.0; +https://sqworm.com/crawler). An alternative string SqwormBot/2.0 may appear in logs from older instances. Behavioral fingerprints include a low variance in request intervals (typically 10–20 ms) and a consistent Accept-Encoding: gzip, deflate header. The bot does not send a Referer header and always includes a Connection: keep-alive header.
📊 Data Usage
The collected data is primarily used for training proprietary AI models hosted by Sqworm Inc., as well as for academic research partnerships. According to Sqworm’s privacy policy, content is stripped of personally identifiable information (PII) before storage, and no raw pages are publicly redistributed. Summarized corpora are made available under a research license to select universities.
⚙️ Rate Limiting Policy
Sqworm is rate-limited because its high request throughput can impact server performance, especially on smaller sites. The policy rationale, as stated in Sqworm’s operator guidelines, is to protect web infrastructure while recognizing that the crawler is legitimate and non-malicious—blocking should only occur after sustained aggressive behavior (e.g., >500 requests in 10 seconds) and never based solely on User-Agent string.
Similar Threats
🛡️
Stop Bots. Save Bandwidth. Protect Revenue.
Boteraser automatically detects and blocks unwanted bots — protecting your site from scrapers, DDoS bursts, and credential stuffing attacks without slowing down real visitors.
✅ Start Free ProtectionSetup takes under a minute · Free trial available
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.