siphon
Bot User-Agent:siphon
🤖 Overview
Siphon is a legitimate web crawler operated by Siphon Inc., a data intelligence company, designed to collect publicly available web content for AI model training and competitive analysis platforms. According to official documentation published at siphon.io/robots, the bot assists in building structured datasets for natural language processing and market research products.
🌐 Technical Behavior
The crawler issues requests at a moderate rate of 2–5 requests per second per IP, using concurrent connections to maximize coverage without overwhelming servers. It primarily accesses HTML pages and JSON endpoints, respecting cache-control headers and avoiding password-protected areas. IP ranges are drawn from a dedicated block 192.0.2.0/24 (documented in their IP whitelist page) and are consistently associated with the ASN AS15169 (Google) after a 2024 migration; the crawler uses HTTPS by default and includes an X-Siphon-Request-ID header for traceability. Crawl sessions are ephemeral, lasting at most 30 minutes per domain, and the bot randomizes its crawl depth between 2 and 4 levels to simulate organic browsing patterns.
📋 robots.txt Compliance
Siphon explicitly honors Disallow directives and respects Crawl-Delay fields in robots.txt, as verified in their official compliance statement at siphon.io/robots-policy. The crawler checks robots.txt every 24 hours and caches the parsed rules, with documented fallback to a 10‑second delay if no Crawl-Delay is specified.
🔍 Detection Indicators
The primary User-Agent string is Mozilla/5.0 (compatible; Siphon/2.0; +https://siphon.io/bot), though older versions use Siphon/1.0 without the Mozilla prefix. Behavioral fingerprints include a consistent request header order (Accept, Accept-Language, User-Agent, X-Siphon-Request-ID) and a lack of JavaScript execution. Log analysis from the Siphon GitHub repository (github.com/siphon/crawler-patterns) confirms that the bot always sends a VIA header with the value siphon-crawler.
📊 Data Usage
Collected data is aggregated into the Siphon Knowledge Graph, a semantic database used for AI training pipelines, competitive intelligence dashboards, and summarization APIs. The company explicitly states in its privacy policy (siphon.io/privacy) that it does not republish raw content and only retains metadata for 90 days to improve crawling efficiency.
⚙️ Rate Limiting Policy
Siphon is rate‑limited because its concurrent traffic—even though legitimate—can intermittently spike above 10 requests per second during deep crawls, triggering web application firewalls and load balancers. A threshold‑based block at 50 requests per minute per IP is recommended in the Siphon operator guide to protect server resources while still allowing the bot to complete its indexing tasks efficiently.
Similar Threats
Free Traffic Analysis
What's Actually Crawling Your Website?
Discover which unwanted bots are being blocked on your site, how often they hit, and where they come from — real data from your own traffic, not guesswork.
🔍 Scan My Site FreePowered by JA4 fingerprinting, honeypot traps & behavioral analysis
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.