CloudVertexBot

Bot User-Agent: cloudvertexbot

🤖 Overview

CloudVertexBot is a web crawler operated by CloudVertex Technologies Inc., a company specializing in cloud-based machine learning and data aggregation. Its primary purpose is to collect publicly available web content to train and improve CloudVertex's proprietary large language models and AI services, including the VertexAI platform. The bot was first documented in early 2024 and has been observed in web server logs since then, according to official statements at cloudvertex.com/bot-policy and their GitHub repository.

🌐 Technical Behavior

CloudVertexBot initiates HTTP GET requests with a default crawl rate of up to 10 requests per second, adjusting based on server response times and the Crawl-Delay value in robots.txt. It uses both HTTP/1.1 and HTTP/2 protocols, supports gzip and deflate compression, and includes an Accept: text/html, application/xhtml+xml header. IP addresses are drawn from the dedicated range 204.15.20.0/24 registered under ASN 39418, as listed in their official IP whitelist at cloudvertex.com/asn. The bot follows a breadth-first traversal strategy and respects Disallow directives, as verified by network analysis published on their developer blog.

📋 robots.txt Compliance

CloudVertexBot fully honors robots.txt directives when a User-agent: CloudVertexBot line is present, including specific Disallow paths and the root exclusion. It also respects the Crawl-Delay directive, pacing requests accordingly. However, according to their documentation at github.com/cloudvertex/crawler-spec, the bot may ignore malformed or ambiguous rules and does not cache robots.txt indefinitely, rechecking every 24 hours.

🔍 Detection Indicators

The primary User-Agent string is CloudVertexBot/1.0 (compatible; CloudVertex Crawler; +https://cloudvertex.com/bot). Additional fingerprinting headers include a From field set to [email protected] and a User-Agent field that may include version suffixes like CloudVertexBot/2.0. Behavioral patterns include sequential, low-variance request intervals and an absence of JavaScript rendering; the bot never submits forms or executes client-side scripts.

📊 Data Usage

All collected content is used to train CloudVertex’s large language models, particularly for natural language understanding, summarization, and code generation tasks. The data also feeds the VertexAI search index, which powers AI-driven responses in CloudVertex’s enterprise products. CloudVertex states that personal data is proactively filtered at ingestion, and they comply with the Do Not Track header where applicable per their privacy policy.

⚙️ Rate Limiting Policy

Despite its legitimate nature, CloudVertexBot’s high request volume—often exceeding 100 requests per minute on busy sites—can degrade server performance. Rate limiting is therefore recommended as a protective measure; typical thresholds like 60 requests per minute per IP are advised by web administrators to balance availability while still allowing the bot to index essential content.

Free Traffic Analysis

What's Actually Crawling Your Website?

Discover which unwanted bots are being blocked on your site, how often they hit, and where they come from — real data from your own traffic, not guesswork.

🔍 Scan My Site Free

Powered by JA4 fingerprinting, honeypot traps & behavioral analysis

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.