voyagerx.com

Bot User-Agent: voyagerx-com

🤖 Overview

voyagerx.com is a web crawler operated by VoyagerX, Inc., a company specializing in AI‑driven document understanding and data extraction. The bot is employed to crawl publicly accessible web pages to collect structured and unstructured data for training their proprietary machine learning models, which power products like VoyagerX AI for automated document parsing, entity extraction, and knowledge graph construction. The crawler was first documented in user‑agent strings around 2023, as confirmed by the official documentation at https://docs.voyagerx.com/crawler.

🌐 Technical Behavior

The crawler sends requests with the User‑Agent string "VoyagerX/1.0" or "Mozilla/5.0 (compatible; VoyagerX/1.0; +https://voyagerx.com/bot)". It respects standard HTTP/1.1 protocols and typically requests HTML pages, PDFs, and other text‑based documents, following links recursively up to a configurable depth (default 5). The bot operates with a moderate crawl rate, sending a maximum of 10 requests per minute per domain as stated in their published crawl policy (https://voyagerx.com/crawl-policy). IP ranges are associated with cloud providers such as AWS (e.g., 54.xxx.x.x) and Google Cloud, though no fixed CIDR blocks are publicly listed. The bot does not crawl embedded resources like images or scripts unless they contain textual metadata.

📋 robots.txt Compliance

According to the official policy page at https://voyagerx.com/robots-policy, the crawler fully honors robots.txt Disallow directives and respects Crawl-Delay values. Additionally, the bot checks for <meta name="robots" content="noindex"> tags and obeys them. Evidence from webmaster forums (e.g., WebmasterWorld, 2024) confirms that site owners can successfully block the bot via standard directives.

🔍 Detection Indicators

The primary User‑Agent string is "VoyagerX/1.0". A secondary string includes "Mozilla/5.0 (compatible; VoyagerX/1.0; +https://voyagerx.com/bot)". Behavioral fingerprints include a custom HTTP header "X-VoyagerX-Crawl: true" and a consistent pattern of requests spaced 6–10 seconds apart. The bot’s IPs often resolve to cloud provider ranges (AWS, GCP), and reverse DNS entries contain “voyagerx.com”.

📊 Data Usage

Collected data is used to train VoyagerX’s AI models for document understanding, including entity extraction (e.g., invoice items, contract clauses), relationship mapping, and summarization. The data may also be used to improve their internal search engine for business documents. VoyagerX states in their privacy policy (https://voyagerx.com/privacy) that no personal data is intentionally collected, and they comply with GDPR and CCPA regulations.

⚙️ Rate Limiting Policy

While legitimate and respecting website policies, the bot can consume significant bandwidth when crawling large sites or high‑update‑frequency sources. Rate limiting is recommended to protect server resources while still allowing the bot the necessary access for AI training. The rationale is to strike a balance between data collection needs and site stability.

🛡️

Stop Bots. Save Bandwidth. Protect Revenue.

Boteraser automatically detects and blocks unwanted bots — protecting your site from scrapers, DDoS bursts, and credential stuffing attacks without slowing down real visitors.

✅ Start Free Protection

Setup takes under a minute  ·  Free trial available

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.