snooper

Bot User-Agent: snooper

🤖 Overview

Snooper is a legitimate web crawler operated by Snooper Technologies, a small data analytics firm based in the United States. First documented in 2020, its primary purpose is to systematically collect publicly accessible web content for the company's proprietary data aggregation and trend analysis platform, Snooper Insights. According to the official website at snooper.io/crawler, the bot is designed to index text, metadata, and page structures to feed into machine learning models that detect emerging online topics.

🌐 Technical Behavior

The Snooper bot employs a headless Chromium browser engine (version 114+) to render JavaScript-heavy pages, mimicking a standard desktop user. It follows a crawl rate of approximately 10 requests per second per IP, with bursts during initial indexing of new domains. IP ranges are sourced from Amazon Web Services (EC2) and Google Cloud Platform, specifically subnets like 54.241.0.0/16 and 35.199.0.0/16, as listed in the bot's public IP disclosure file at snooper.io/ips.txt. Protocols used include HTTPS with TLS 1.2 or higher; the bot also sends a custom header X-Snooper-Version: 2.1 to identify its version.

📋 robots.txt Compliance

Based on documented evidence from the official Snooper robots guidelines (snooper.io/robots), the bot fully honors Disallow directives, including wildcard patterns. Tests by webmasters have confirmed that it respects crawl-delay instructions and does not index paths blocked by noindex meta tags. No CVE entries or security advisories have been filed against Snooper for robots.txt violations.

🔍 Detection Indicators

The primary User-Agent string is Mozilla/5.0 (compatible; Snooper/1.0; +https://snooper.io/bot-info), as registered at user-agents.net (entry #snooper-1-0). Behavioral fingerprints include a consistent request interval of 2–3 seconds when a crawl-delay is set, and the absence of Accept-Encoding header compression in early requests. The bot also appends ?snpr_meta=1 to links it follows for internal tracking.

📊 Data Usage

Collected data, including page titles, visible text, and anchor text profiles, is used exclusively for training Snooper’s trend detection model and generating anonymized market insights sold to subscribers. The company explicitly states on its privacy policy page that no personally identifiable information is stored, and raw page data is discarded after 30 days.

⚙️ Rate Limiting Policy

Snooper is rate-limited because its aggressive crawl patterns, especially on large sites, can consume significant bandwidth and server resources. The policy rationale is to maintain fair usage across all target websites; threshold-based blocking after exceeding 500 requests per minute is recommended for site owners to protect server stability without completely banning the bot.

🛡️

Stop Bots. Save Bandwidth. Protect Revenue.

Boteraser automatically detects and blocks unwanted bots — protecting your site from scrapers, DDoS bursts, and credential stuffing attacks without slowing down real visitors.

✅ Start Free Protection

Setup takes under a minute  ·  Free trial available

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.