hyperix

Bot User-Agent: hyperix

🤖 Overview

Hyperix is a web crawler operated by Hyperix Ltd., a data intelligence company headquartered in London. First publicly documented in a 2022 blog post on the company’s website (https://hyperix.io/about), the bot is designed to collect publicly available web content for the Hyperix Data Platform, which provides SEO analytics, competitive intelligence, and AI model training datasets. According to the official Hyperix documentation, the crawler indexes over 1 billion pages per month to power their “Web Insights” product.

🌐 Technical Behavior

The Hyperix bot uses a distributed crawling architecture deployed on Amazon Web Services and Google Cloud Platform, with IP ranges documented in the company’s public IP list at https://hyperix.io/ips. It issues requests at a rate of 10–20 per second per IP, but can burst to 50 during deep crawl sessions. The crawler follows HTTP redirects and respects robots.txt directives for crawl-delay, though it does not support the Crawl-Delay directive itself – instead it uses a fixed rate. It sends a User‑Agent header, an Accept-Language header set to “en-US,en;q=0.9”, and a Referer field set to its own website. The bot also checks for sitemap.xml files to prioritize pages, as noted in the official technical documentation.

📋 robots.txt Compliance

Per the Hyperix official guidelines at https://hyperix.io/robots-policy, the bot unconditionally honors Disallow directives and will not crawl any path explicitly excluded. The documentation states that the bot also respects the Crawl-Delay directive when present in robots.txt, despite not using it for its own rate limiting. It will not follow Nofollow meta tags or rel=”nofollow” links, but will respect Noindex directives at the page level.

🔍 Detection Indicators

The primary identifying string is “HyperixBot/1.0” with the full User‑Agent “Mozilla/5.0 (compatible; HyperixBot/1.0; +https://hyperix.io/bot)”. It also uses a secondary agent “HyperixCollector/2.0” for its analytics pipeline. Behavioral fingerprints include a consistent request rate of exactly 1500ms between requests and a lack of cookies or JavaScript execution. The bot always sends a custom header “X-Hyperix-Request: true”.

📊 Data Usage

Collected data is ingested into the Hyperix Data Platform, where it is used to train proprietary AI models for SEO content recommendations, market trend analysis, and real‑time competitor monitoring. The platform also exposes aggregated data to subscribers through an API, as described in the Hyperix service terms (https://hyperix.io/terms). No raw content is sold; instead, derived insights are provided to paying customers.

⚙️ Rate Limiting Policy

Hyperix is rate‑limited on many web applications because its distributed architecture can produce sustained high request volumes that degrade server performance. Standard rate‑limiting thresholds (e.g., 100 requests per minute per IP) are considered fair use by the company, which recommends contacting them at [email protected] for custom crawl arrangements if higher rates are needed.

Free Traffic Analysis

What's Actually Crawling Your Website?

Discover which unwanted bots are being blocked on your site, how often they hit, and where they come from — real data from your own traffic, not guesswork.

🔍 Scan My Site Free

Powered by JA4 fingerprinting, honeypot traps & behavioral analysis

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.