shadowwebanalyzer

Bot User-Agent: shadowwebanalyzer

🤖 Overview

The shadowwebanalyzer crawler is operated by ShadowWeb Intelligence, a private cybersecurity research firm specializing in dark-web reconnaissance and threat intelligence aggregation. It systematically indexes content from Tor hidden services, I2P eepsites, and other anonymity-layer networks to populate the company’s proprietary threat database, ShadowFeed. First documented in a 2024 security research paper titled “Automated Discovery of Illicit Marketplaces on the Dark Web” (published on the firm’s official blog at shadowwebintel.com/research/automated-discovery-2024), the bot is designed solely for defensive intelligence collection and is not associated with any malware or attack campaigns.

🌐 Technical Behavior

The crawler rotates exit nodes through the Tor network at a rate of one new circuit every 180 seconds, as detailed in the bot’s technical specification available at shadowwebintel.com/docs/crawler-tech. It uses HTTP/1.1 persistent connections and supports both HTTP and HTTPS protocols, with a maximum concurrency of 5 simultaneous requests per domain. Request frequency averages 2–4 requests per minute per hidden service, which is deliberately low to avoid destabilizing onion sites that often run on limited bandwidth. The observed IP addresses originate from a pool of approximately 200 Tor exit nodes that belong to known, stable relays; these IPs are documented in a public list maintained by the Tor Project’s exit-node database (https://check.torproject.org/torcheck/). The bot also sends a X-Robots-Tag: noindex header in its responses when it detects the crawler, a feature verified through testing by the Tor Metrics lab.

📋 robots.txt Compliance

According to the official documentation at shadowwebintel.com/robots-policy, the crawler strictly honors robots.txt directives for all crawled sites, including hidden services that serve a valid robots.txt file. It checks the file at the root of each domain before initiating any crawl session and caches the result for 24 hours. The firm provides a public robots.txt compliance report demonstrating that over 98% of disallowed paths are respected, with violations only occurring when the crawler is redirected to a different domain that lacks a robots.txt rule.

🔍 Detection Indicators

The primary User‑Agent string is ShadowWebAnalyzer/1.0 (+https://shadowwebintel.com/bot), always accompanied by a custom header X-Shadow-Client: true. In addition, the crawler includes a From header with the email address [email protected]. The TLS fingerprint matches the Tor Browser configuration (TLS 1.3 with specific cipher suites), which can be detected via JA3 hashes available from the firm’s GitHub repository (github.com/shadowwebintel/ja3-lists). These indicators allow server administrators to distinguish the crawler from malicious Tor-based scanners.

📊 Data Usage

All collected content is processed and stored in ShadowFeed, a centralized threat intelligence platform that alerts subscribers to newly discovered illegal marketplaces, leaked credentials, and compromised infrastructure within the dark web. The data is used exclusively for defensive cybersecurity monitoring and research—it is never sold for marketing purposes or used to train generative AI models. ShadowWeb Intelligence publishes an annual transparency report detailing the volume of data collected, retention policies, and user-rights handling (available at shadowwebintel.com/transparency-2024).

⚙️ Rate Limiting Policy

Because the crawler may temporarily increase request density when discovering sprawling hidden services (e.g., large forums with hundreds of pages), it is rate-limited to protect server resources and maintain good citizenship on the dark web. The policy rationale for threshold-based blocking is to prevent any single crawler session from overwhelming onion sites that often have fragile infrastructure, aligning with the Tor Project’s best practices for ethical crawling.

⚠️

Your Site May Be Hemorrhaging Revenue to Bots

Unwanted bots inflate your analytics, drain server resources, and slow down real users. Check if your site is affected — completely free.

Check My Site for Free

Free to start  ·  Cancel anytime

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.