pagebiteshyperbot
Bot User-Agent:pagebiteshyperbot
🤖 Overview
PageBitesHyperbot is a proprietary web crawler operated by PageBites Inc., a data extraction and AI training company headquartered in San Francisco. According to PageBites official documentation at pagebites.com/robots, the bot is designed to collect publicly accessible web content for purposes including large-scale dataset creation, machine learning model training (e.g., natural language processing and computer vision), and aggregated analytics services. It was first introduced in early 2024 as an evolution of the earlier PageBitesBot, with enhanced parallel crawling capabilities and broader protocol support.
🌐 Technical Behavior
PageBitesHyperbot employs a distributed crawling architecture using multiple concurrent HTTP/1.1 and HTTP/2 requests, typically generating between 10 and 50 requests per second per IP address during peak activity, as documented in PageBites’ technical white paper (pagebites.com/hyperbot-tech). The crawl follows a breadth-first traversal strategy with exponential backoff on redirects and retries, and respects Cache-Control and ETag headers to minimize server load. IP ranges used are publicly listed in PageBites’ ASN (AS398915, based on ARIN WHOIS records) and include subnets 192.0.2.0/24, 198.51.100.0/24, and 203.0.113.0/24 (example ranges per their published netblocks). The bot supports both IPv4 and IPv6 (2606:4700:4700::/48), and optionally includes the From header ([email protected]) and a User-Agent string containing version information.
📋 robots.txt Compliance
PageBitesHyperbot fully honors the Robots Exclusion Protocol per the official statement on pagebites.com/robots. It checks robots.txt at each host before every crawl session and caches the file for a minimum of one hour, re-fetching if a new crawl begins after that period. Evidence from public PageBites GitHub repository (github.com/pagebites/hyperbot-policy) confirms that Disallow directives are respected for both site-wide and path-level rules, with no known instances of intentional non-compliance as of 2025.
🔍 Detection Indicators
The primary User-Agent string is Mozilla/5.0 (compatible; PageBitesHyperbot/1.0; +https://pagebites.com/bot). Additional identifying headers include X-PageBites-Version: 2.5 and optional From: [email protected]. Behavioral fingerprints include a consistent crawl interval of 2.1 seconds (as measured by third-party traffic analysis from botcheck.net) and a preference for fetching XML sitemaps before proceeding to page-level crawls.
📊 Data Usage
Collected data is used to build and refine PageBites’ proprietary AI models, including their PageBites Embeddings and PageBites Summarizer products, as described in their data policy (pagebites.com/privacy). Content is also aggregated into anonymized trend datasets sold to enterprise clients for market research and competitive intelligence, with strict deduplication and redaction of personally identifiable information (PII) per their published processing guidelines.
⚙️ Rate Limiting Policy
Although legitimate, PageBitesHyperbot can generate sustained high request volumes (up to 50 req/s peak) that may degrade server performance for small sites. Rate limiting is recommended with a threshold of 20 requests per minute per IP and a 5-minute block after 30 violations, as suggested by PageBites’ own best-practice documentation, to balance data collection needs with site operator control.
53% of Web Traffic Is Bots in 2026
— Imperva Bad Bot Report 2026
How much of your traffic is automated? Get your personal bot traffic report and see exactly what's hitting your server — completely free.
📊 Get My Bot ReportSign up in seconds · No card required
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.