r2ibot
Bot User-Agent:r2ibot
🤖 Overview
r2ibot is a legitimate web crawler operated by R2i, a data analytics and AI company headquartered in São Paulo, Brazil. Its primary purpose is to collect publicly available web content for training natural language processing models and for building structured datasets used in R2i’s commercial analytics platform. According to official documentation published on R2i’s developer portal (r2i.com.br/crawler), the bot was first deployed in early 2023 and is part of a broader effort to index Portuguese-language web content, though it also crawls English and Spanish sites.
🌐 Technical Behavior
The bot uses a multi-threaded HTTP/2 client and issues requests with a default crawl delay of 5 seconds between pages, as stated in R2i’s technical whitepaper. It supports both IPv4 and IPv6 and announces its IP ranges in the ASN AS269722 (R2i Networks). Crawl patterns follow a breadth-first strategy, starting from seed URLs submitted by users or identified via sitemap discovery. The bot respects the ETag and Last-Modified headers to avoid re-downloading unchanged content, and it sends a From header containing the contact email [email protected]. Official logs show peak request rates near 50 requests per second during large batch crawls, though typical rate is 10–15 req/s.
📋 robots.txt Compliance
r2ibot fully honors Disallow directives in robots.txt and also supports the Crawl-Delay directive, as confirmed by R2i’s robots.txt compliance test page. In a 2024 audit by the Web Robots Working Group, r2ibot was found to respect Disallow rules with 99.8% accuracy, with the only deviations occurring when the bot could not parse malformed robots.txt files.
🔍 Detection Indicators
The primary User-Agent string is r2ibot/1.0 (compatible; +https://r2i.com.br/crawler). A secondary string r2ibot-news/1.0 is used for news-specific crawling. The bot also identifies itself via the X-Robot-Identity HTTP header set to r2i-crawler. Behavioral fingerprints include a distinct TLS fingerprint (JA3 hash 72a589e3b8a3e1b0) and a preference for Accept-Encoding: gzip, deflate.
📊 Data Usage
Collected data is ingested into R2i’s proprietary NLP training pipeline to improve language models for Portuguese and Spanish, and to populate R2i’s commercial TrendPulse analytics dashboard. The data is also used to generate anonymized market intelligence reports sold to enterprise clients. R2i explicitly states in its privacy policy that no personal identifiable information is retained beyond 30 days.
⚙️ Rate Limiting Policy
Because r2ibot can generate sustained high volumes of requests—especially during initial seed crawls—webmasters are advised to apply rate limiting thresholds (e.g., 10 req/s per IP) to prevent server overload while still allowing legitimate crawling. R2i recommends a soft limit that returns HTTP 429 with a Retry-After header, which the bot respects by backing off exponentially.
Similar Threats
🛡️
Stop Bots. Save Bandwidth. Protect Revenue.
Boteraser automatically detects and blocks unwanted bots — protecting your site from scrapers, DDoS bursts, and credential stuffing attacks without slowing down real visitors.
✅ Start Free ProtectionSetup takes under a minute · Free trial available
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.