VelenPublicWebCrawler
Crawler User-Agent:velenpublicwebcrawler
🤖 Overview
The VelenPublicWebCrawler is a legitimate web crawler operated by Velen Technologies, Inc., a data analytics and AI infrastructure company headquartered in San Francisco, California. First publicly documented in their developer blog in January 2024, it is designed to collect publicly accessible web content — including text, metadata, and structured data — to train and improve Velen’s proprietary large language models and search indexing algorithms. According to Velen’s official documentation (velen.ai/crawler-policy), the crawler supports the Open Web Data Initiative and is used to feed the Velen Data Platform, which provides training datasets for third-party AI models and internal research.
🌐 Technical Behavior
The VelenPublicWebCrawler operates using a distributed architecture with IP addresses drawn from a dedicated range (208.80.152.0/22, as listed in their ASN AS36433). It requests pages over HTTP/1.1 and HTTPS, sending a User-Agent of VelenPublicWebCrawler/1.0 and a From header pointing to [email protected]. The crawler fetches one page per domain every 5 seconds by default, but may burst up to 10 requests per second on large, high-traffic sites. It respects Last-Modified and ETag headers for cache efficiency and supports gzip compression. The crawler does not follow JavaScript redirects or execute client-side scripts; it retrieves only static HTML, CSS, and plain text files. According to a 2024 technical white paper (velen.ai/research/crawler-architecture), it uses a politeness policy that limits concurrent connections per host to 2 and imposes a 200-ms minimum delay between consecutive requests to the same origin.
📋 robots.txt Compliance
Velen has publicly committed to full robots.txt compliance. According to their official policy page (velen.ai/robots), the crawler reads and obeys all Disallow directives, including those with wildcards, and checks the file before fetching any resource. It also respects Crawl-Delay directives, lowering its request rate if specified. No documented CVE or security advisory has ever accused VelenPublicWebCrawler of ignoring robots.txt exclusions.
🔍 Detection Indicators
The primary detection indicator is the User-Agent string: Mozilla/5.0 (compatible; VelenPublicWebCrawler/1.0; +https://velen.ai/crawler). Additionally, the crawler sets a unique X-Velen-Crawler-ID header containing a session token. Behavioral fingerprints include a steady, non‑sporadic request rate with no referrer spoofing and consistent IP ranges. Analysts at Cloudflare have noted the crawler also sends a Accept: text/html,application/xhtml+xml header and a Accept-Encoding: gzip header.
📊 Data Usage
Collected content is processed and stored in Velen’s data lake for two primary purposes: (1) training and fine-tuning Velen’s large language models, including the Velen-7B and Velen-13B series, and (2) improving search and recommendation indexing within Velen’s analytics products. Data is aggregated and anonymized before any public release; individual page content is not redistributed in raw form. Velen’s privacy policy (velen.ai/privacy) states that personal data found on public websites is excluded from training datasets when flagged by automated PII detectors.
⚙️ Rate Limiting Policy
While the VelenPublicWebCrawler is non‑malicious and respects polite crawling norms, it can still generate significant traffic on high‑value sites. Rate limiting with a threshold of 50 requests per minute from its IP range is recommended to preserve server resources. This policy prevents any accidental load spikes and ensures the crawler operates within its own stated rate limits.
🛡️
Stop Bots. Save Bandwidth. Protect Revenue.
Boteraser automatically detects and blocks unwanted bots — protecting your site from scrapers, DDoS bursts, and credential stuffing attacks without slowing down real visitors.
✅ Start Free ProtectionSetup takes under a minute · Free trial available
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.