Skip to main content

Boteraser | Website and Server Security Solutions

webinator-wbi

Bot User-Agent: webinator-wbi

🤖 Overview

webinator-wbi is a web crawler operated by Thunderstone Software, the developer of the Webinator enterprise search and indexing platform, first released in the late 1990s. Its purpose is to crawl websites, intranets, and document repositories to build a searchable index for internal or external use, as documented on Thunderstone's official site at https://www.thunderstone.com/texis/webinator/. The bot has also been referenced in academic research on web crawling and is not associated with AI training or third-party data harvesting.

🌐 Technical Behavior

The crawler issues sequential HTTP GET requests, obeying a configurable crawl delay that defaults to several seconds between requests. It respects robots.txt, handles sitemap files, and obeys meta robots tags and the X-Robots-Tag HTTP header. IP addresses originate from Thunderstone's cloud infrastructure (typically within their ASN) when using hosted service, or from the customer's own IP range when self-hosted. The bot supports conditional GET via If-Modified-Since and ETag headers to reduce bandwidth usage, and can follow links automatically. Newer versions can be configured to crawl JavaScript-rendered content using headless browsing. The default request frequency is moderate, but administrators may increase it significantly for large sites, making rate-limiting advisable.

📋 robots.txt Compliance

Thunderstone explicitly states that webinator-wbi fully honors all robots.txt directives, including Disallow, Allow, and Crawl-delay. It also respects the Robots Exclusion Protocol standard and the X-Robots-Tag header for non-HTML content. Compliance is hard-coded into the Webinator software, and site operators can verify this through the official manual.

🔍 Detection Indicators

The primary User-Agent string is webinator-wbi, often with a version suffix such as webinator-wbi/4.0 or webinator-wbi/5.0. Common formats include Mozilla/5.0 (compatible; webinator-wbi; +http://www.thunderstone.com/texis/webinator/). Some installations include the administrator's email in a From header. The bot supports Accept-Encoding (gzip) and Connection: keep-alive, but does not send unusual or spoofed headers. The presence of the token "webinator-wbi" is definitive for identification.

📊 Data Usage

Collected content is used solely to populate Webinator's search index, providing full-text search, faceted navigation, and relevance ranking for websites and intranets. No data is repurposed for AI training, machine learning, or sold to third parties. Indexed data remains under the control of the organization running Webinator; analytics are limited to search performance metrics internal to the platform.

⚙️ Rate Limiting Policy

Rate limiting of webinator-wbi is recommended because its crawl speed is entirely configurable by the operator, potentially causing high request volumes that impact server performance. Threshold-based blocking ensures fair resource allocation while still permitting legitimate indexing activity, aligning with standard webmaster best practices for non-malicious crawlers.

🛡️

Stop Bots. Save Bandwidth. Protect Revenue.

Boteraser automatically detects and blocks unwanted bots — protecting your site from scrapers, DDoS bursts, and credential stuffing attacks without slowing down real visitors.

✅ Start Free Protection

Setup takes under a minute  ·  Free trial available

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.