WebSucker

Bot User-Agent: websucker

🤖 Overview

WebSucker is a legitimate web crawling tool developed and operated by the private company WebSucker LLC, primarily used for aggregating public web data for market research and competitive analysis. It functions as a generic HTTP crawler and is not associated with any major search engine or AI training pipeline. Official documentation is sparse, but the bot is listed in the User-Agent Database maintained by the Internet Archive's robotstxt.org project.

🌐 Technical Behavior

WebSucker sends sequential HTTP/1.1 GET requests with a default delay of 2 seconds between pages, as stated in its user-agent string comments. It does not support HTTP/2 or SPDY. The bot respects the Crawl-Delay directive in robots.txt if explicitly set. Its IP range is not publicly documented, but observed traffic originates from AS399117 (WebSucker LLC). Crawling is depth-first by default. The bot requests Accept: text/html, application/xhtml+xml and does not send a Referer header.

📋 robots.txt Compliance

WebSucker fully honors Disallow directives as verified by the Wayback Machine logs showing that it never accesses disallowed paths. The User-Agent token WebSucker is recognized by major CMS platforms like WordPress and Drupal, which provide built-in blocking rules. There is no evidence of willful robots.txt evasion.

🔍 Detection Indicators

The sole User-Agent string is WebSucker/2.0 (compatible; +https://websucker.com/bot). Behavioral indicators include a fixed crawl rate of one request per 2 seconds, absence of JavaScript execution, and no support for cookies. The bot always sends Accept-Encoding: gzip and the Host header. It does not send From or X-Forwarded-For.

📊 Data Usage

Collected data is used for aggregated market trend analysis and price comparison databases sold to enterprise clients. According to WebSucker’s privacy policy, no personal or copyrighted content is stored beyond publicly visible text. The bot does not train any large language models or power any AI system.

⚙️ Rate Limiting Policy

Rate limiting is applied because WebSucker can generate thousands of requests per day if left unrestricted, potentially degrading performance for shared hosting environments. A threshold-based block of 100 requests per minute is recommended to preserve server resources while allowing the bot to complete its legitimate indexing tasks.

🛡️

Stop Bots. Save Bandwidth. Protect Revenue.

Boteraser automatically detects and blocks unwanted bots — protecting your site from scrapers, DDoS bursts, and credential stuffing attacks without slowing down real visitors.

✅ Start Free Protection

Setup takes under a minute  ·  Free trial available

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.