Rainbot
Bot User-Agent:rainbot
🤖 Overview
Rainbot is a web crawler operated by Rainbird (rainbird.ai), a company specializing in knowledge automation and AI-driven decision intelligence. It collects publicly available web data to train Rainbird's knowledge models and improve the accuracy of its expert system platform, which is used by enterprises for complex decision-making processes. Rainbot was first documented in 2020 and is officially designated as a legitimate, rate-limited crawler in Rainbird’s developer documentation at docs.rainbird.ai.
🌐 Technical Behavior
Rainbot performs HTTP/1.1 and HTTP/2 requests with a configurable crawl delay defaulting to 10 seconds between consecutive requests, as specified in its robot‐exclusion policy. It targets text‑based content (HTML, JSON, XML) and avoids binary files. The crawler originates from IP ranges belonging to Amazon Web Services (AWS) and Google Cloud Platform (GCP), specifically subnets 52.6.0.0/15 and 35.190.0.0/16, according to Rainbird’s published network policy. Rainbot supports conditional GET requests using ETags and Last‑Modified headers to minimize bandwidth. It adheres to the crawl‑delay directive and respects noindex meta tags.
📋 robots.txt Compliance
Rainbird’s official site states that Rainbot fully honors robots.txt Disallow directives. Any site can block it entirely by adding User‑agent: Rainbot followed by Disallow: /. Community reports on webmaster forums confirm that Rainbot does not ignore these rules, making it a well‑behaved, non‑aggressive crawler.
🔍 Detection Indicators
The primary User‑Agent string is "Rainbot/1.0 (Rainbird; +https://rainbird.ai/rainbot)". Secondary strings include "Rainbird-Crawler" with version suffixes. Rainbot also sends a custom HTTP header X‑Rainbird‑Crawler: 1 for unambiguous identification. The From header contains the contact address [email protected].
📊 Data Usage
Collected data is used exclusively to train Rainbird’s knowledge automation AI models that power enterprise decision‑logic systems. The crawled content helps the platform understand domain‑specific relationships and hierarchies. Rainbird does not use the data for advertising or profiling, and retains it only as long as necessary for model improvement, as per its privacy policy at rainbird.ai/privacy.
⚙️ Rate Limiting Policy
Rainbot is rate‑limited because its default concurrency (up to 8 parallel connections) can overload shared hosting environments. Threshold‑based rate limiting triggers after 500 requests per minute, imposing a 48‑hour ban to prevent resource exhaustion while still allowing legitimate indexing.
Similar Threats
🛡️
Stop Bots. Save Bandwidth. Protect Revenue.
Boteraser automatically detects and blocks unwanted bots — protecting your site from scrapers, DDoS bursts, and credential stuffing attacks without slowing down real visitors.
✅ Start Free ProtectionSetup takes under a minute · Free trial available
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.