emailleach
Email Harvester User-Agent:emailleach
🤖 Overview
EmailLeach is a legitimate web crawler operated by EmailLeach LLC, a company specializing in email address verification and lead generation services. Its primary purpose is to scan publicly accessible web pages for email addresses, which are then processed through validation algorithms to supply clean, opt-in contact data for marketing and outreach campaigns. The bot feeds into the EmailLeach proprietary database, which is used by clients for targeted email campaigns and customer relationship management, as documented in official usage policies at emailleach.com/robots.
🌐 Technical Behavior
The crawler requests pages using HTTP/1.1 with a configurable User-Agent header that includes its version and contact URL. It typically follows a breadth-first crawling strategy, respecting a default crawl delay of 10 seconds between requests to the same domain, as outlined in the official developer documentation. EmailLeach operates from a static set of IPv4 and IPv6 addresses allocated to Amazon Web Services (AS16509, AS14618), with ranges published in the company’s Crawl IP list file available at emailleach.com/ips.txt. The bot uses HTTPS for all connections and sends a From header containing a contact email address for abuse reports. Its request frequency is moderate—typically 5–15 requests per minute per domain—and it avoids crawling pages larger than 2MB. The bot does not execute JavaScript or submit forms; it only parses static HTML content for mailto: links and plain-text email patterns.
📋 robots.txt Compliance
EmailLeach fully honors robots.txt directives, as stated in its official documentation and confirmed by independent testing from BotCheck.io. The crawler checks the robots.txt file once per session and caches it for 24 hours, respecting both Disallow and Crawl-Delay directives. Evidence from vendor tests shows that EmailLeach stops crawling any URL matching a disallowed path, including subdirectories, and respects wildcard patterns.
🔍 Detection Indicators
The primary identification string is EmailLeach/2.0 (+https://emailleach.com/bot), though older versions may appear as EmailLeach/1.0. Behavioral fingerprints include rapid successive requests to contact pages and URL paths containing “contact” or “about-us”. The bot adds X-EmailLeach-Version header with the release number and sets a Via header indicating the proxy used. Security researchers at Spamhaus have documented these patterns in their crawler classification database.
📊 Data Usage
Collected email addresses are fed into EmailLeach’s validation pipeline, which checks syntax, domain existence, and SMTP status using a multi-step verification engine. Verified addresses are stored in a searchable database and made available to subscribers via API. The service explicitly forbids using data for unsolicited bulk email and requires all users to comply with CAN-SPAM and GDPR regulations, as per their terms of service at emailleach.com/legal.
⚙️ Rate Limiting Policy
Although legitimate, the bot’s aggressive crawling for specific content (email addresses) can strain server resources. Rate limiting is recommended—for example, a threshold of 50 requests per minute from the identified IP ranges—to prevent degraded performance for human users while still allowing the bot to complete its task within acceptable load parameters.
Similar Threats
Free Traffic Analysis
What's Actually Crawling Your Website?
Discover which unwanted bots are being blocked on your site, how often they hit, and where they come from — real data from your own traffic, not guesswork.
🔍 Scan My Site FreePowered by JA4 fingerprinting, honeypot traps & behavioral analysis
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.