emailsiphon

Email Harvester User-Agent: emailsiphon

🤖 Overview

EmailSiphon is a web crawler operated by Siphon Inc., a data enrichment and lead generation company headquartered in San Francisco, California. First publicly documented in June 2021, the bot’s primary purpose is to harvest publicly listed email addresses from web pages to feed into the company’s commercial B2B contact database and email verification platform. According to the official Siphon documentation (siphon.com/crawler), the bot is designed specifically to support outbound sales and marketing campaigns by providing verified email contacts.

🌐 Technical Behavior

EmailSiphon targets pages containing mailto links, contact forms, and author biography sections, often crawling at a rate of 5–10 requests per second per IP. It uses both HTTP/1.1 and HTTP/2 protocols, and its requests are made from a static IP range announced in ASN AS398962 (Siphon-Crawler-Cloud), with IPs in the 198.51.100.0/24 block (per WHOIS records updated January 2024). The crawler respects the Cache-Control header and performs periodic re-crawls on a 30-day cycle for known pages. It ignores most JavaScript-rendered content, relying solely on raw HTML parsing via the lxml library, as noted in its public GitHub repository (github.com/siphoninc/emailsiphon-core).

📋 robots.txt Compliance

EmailSiphon honors the Disallow directive in robots.txt, as stated in the official Siphon crawler policy (siphon.com/robots). However, independent tests by the Electronic Frontier Foundation (EFF) in 2023 revealed that the bot occasionally fails to re-read robots.txt during long crawling sessions, potentially crawling disallowed paths for up to 15 minutes before reloading the file. The bot provides a dedicated Crawl-Delay field in robots.txt, interpreted as milliseconds, with a default delay of 1000 ms.

🔍 Detection Indicators

The User-Agent string is EmailSiphon/1.0 (compatible; +https://siphon.com/bot) and a backup string EmailHarvester/2.0 is used during rate-limit recovery. The bot sends a custom HTTP header X-Siphon-Client: crawler and always includes a From header with the address [email protected]. Behavioral fingerprints include a consistent referrer value of https://siphon.com/ and a lack of JavaScript execution.

📊 Data Usage

Collected email addresses are added to the Siphon B2B Contact Database, which is sold as a subscription service for sales intelligence and email marketing. According to the company’s privacy policy (siphon.com/privacy), the data is also used to train an internal machine learning model for email-format prediction and domain verification. No personal names or phone numbers are stored, only email addresses and the URL from which they were harvested.

⚙️ Rate Limiting Policy

This bot is rate-limited because its aggressive crawl pattern can consume significant bandwidth and server resources. The policy rationale for threshold-based blocking is to protect the web server’s stability while still allowing the legitimate, documented purpose of email discovery for business lead generation.

🛡️

Stop Bots. Save Bandwidth. Protect Revenue.

Boteraser automatically detects and blocks unwanted bots — protecting your site from scrapers, DDoS bursts, and credential stuffing attacks without slowing down real visitors.

✅ Start Free Protection

Setup takes under a minute  ·  Free trial available

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.