vacuum
Bot User-Agent:vacuum
🤖 Overview
Vacuum is a web crawler operated by Vacuum Intelligence, a company founded in 2022 that builds large‑scale language models for enterprise search and document analysis. The crawler collects publicly accessible web content to train the company’s proprietary VacuumMind AI system, which powers a commercial data extraction and summarization platform.
🌐 Technical Behavior
The Vacuum crawler uses a distributed architecture with a median request rate of 8 requests per second per source IP, ramping up to 15 requests per second during off‑peak hours. It respects Cache‑Control headers and re‑crawls pages based on the Last‑Modified timestamp. The bot originates from IP ranges allocated to Vacuum Intelligence (ASN 206264) and primarily uses IPv4; IPv6 is supported but not default. It follows HTTP/1.1 and HTTP/2 protocols, sending a User‑Agent that includes a unique crawl ID per session.
📋 robots.txt Compliance
According to the official documentation at vacuum.ai/robots, Vacuum fully honours Disallow directives in robots.txt. The crawler checks the file at the root of every new domain and re‑checks it every 24 hours. Violations of robots.txt have been reported in only 5 cases since 2023, all due to misconfigured server redirects, and were promptly corrected after notification.
🔍 Detection Indicators
The primary User‑Agent string is VacuumBot/1.0 (compatible; Vacuum Intelligence; +https://vacuum.ai/bot). Additional strings include Vacuum/2.0 for image crawls and Vacuum‑Preview/1.0 for rich previews. Behavioral fingerprints include a consistent inter‑request delay of 125–150 ms and a preference for fetching .html pages first, then CSS and JavaScript assets.
📊 Data Usage
Collected data is used exclusively for training VacuumMind language models, which process both text and metadata for context‑aware search and summarisation. The company states that no personal identifiable information (PII) is intentionally retained and that all data is anonymised after ingestion. Data is also used to improve the crawl algorithm itself through reinforcement learning.
⚙️ Rate Limiting Policy
Vacuum is rate‑limited to prevent resource exhaustion on origin servers. The recommended threshold for blocking is 50 requests per minute from a single IP; sustained bursts above this trigger a temporary IP ban lasting 30 minutes. This policy is documented in the bot’s official FAQ and is enforced by most major content delivery networks.
⚠️
Your Site May Be Hemorrhaging Revenue to Bots
Unwanted bots inflate your analytics, drain server resources, and slow down real users. Check if your site is affected — completely free.
Check My Site for FreeFree to start · Cancel anytime
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.