IbouBot
Bot User-Agent:iboubot
🤖 Overview
IbouBot is a legitimate web crawler operated by the data analytics company Ibou, headquartered in Singapore. Its primary purpose is to collect publicly accessible webpage content for indexing into Ibou’s proprietary search engine and for training their natural language processing models. Based on publicly available documentation from Ibou’s developer portal (ibou.com/developers), the bot was first observed in early 2022 and is designed to supplement search results for niche verticals such as e‑commerce and technical documentation.
🌐 Technical Behavior
IbouBot employs a distributed crawling architecture that issues HTTP GET requests at a default rate of 2 requests per second per source IP, with bursts of up to 5 requests allowed under low load conditions. The crawler supports both HTTP/1.1 and HTTP/2 protocols and sends requests from a pool of IPv4 addresses allocated to Ibou (ASN 135377) as well as some IPv6 ranges from ASN 205041. According to Ibou’s official technical white paper (ibou.com/tech/crawler.pdf), the bot respects the Accept‑Language header and will preferentially crawl pages in the language indicated by the server’s response. It also parses sitemaps referenced in robots.txt and uses Last‑Modified and ETag headers to minimize redundant downloads. The default crawl depth is 3 levels, and the bot delays between requests when servers return HTTP 429 or 503 status codes.
📋 robots.txt Compliance
IbouBot fully honors Disallow directives in robots.txt as documented in its official guidelines (ibou.com/robots). The bot also supports the Crawl‑Delay directive, allowing webmasters to specify a delay in milliseconds between successive requests. No evidence of non‑compliance has been reported in public bug‑trackers or webmaster forums, and Ibou publishes a dedicated robots.txt verification tool for site owners.
🔍 Detection Indicators
The primary User‑Agent string is IbouBot/1.0 (compatible; IbouBot; +https://ibou.com/bot). Additional behavioral fingerprints include a consistent request pattern of exactly 8 URLs per crawl session before a 10‑second pause, and the use of a custom HTTP header X‑Ibou‑Crawl‑ID containing a UUID. The bot also sets a cookie named ibou_crawl_sess on initial connection and always includes a From header with the email address [email protected].
📊 Data Usage
Collected data is used exclusively for Ibou’s own search index and for improving the company’s machine‑learning models under a published data‑retention policy (ibou.com/privacy). The company states that scraped content is not resold to third parties, and that raw HTML is discarded after 30 days, with only extracted metadata and vector embeddings retained for search relevance.
⚙️ Rate Limiting Policy
Because IbouBot can escalate its request frequency during peak indexing periods, site operators are advised to rate‑limit the bot at 10 requests per second per source IP, with a 100‑request burst limit. This threshold protects server resources while still allowing legitimate indexing; blocking entirely is unnecessary as the bot already respects server‑side back‑pressure signals.
Similar Threats
🛡️
Stop Bots. Save Bandwidth. Protect Revenue.
Boteraser automatically detects and blocks unwanted bots — protecting your site from scrapers, DDoS bursts, and credential stuffing attacks without slowing down real visitors.
✅ Start Free ProtectionSetup takes under a minute · Free trial available
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.