webinator-wbi
webinator-wbi is a web crawler operated by Thunderstone Software, the developer of the Webinator enterprise search and indexing platform, first released in the late 1990s. Its purpose is to crawl websites, intranets, and document repositories to build a searchable index for internal or external use, as documented on Thunderstone's official site at https://www.thunderstone.com/texis/webinator/. The bot has also been referenced in academic research on web crawling and is not associated with AI training or third-party data harvesting.
The crawler issues sequential HTTP GET requests, obeying a configurable crawl delay that defaults to several seconds between requests. It respects robots.txt, handles sitemap files, and obeys meta robots tags and the X-Robots-Tag HTTP header. IP addresses originate from Thunderstone's cloud infrastructure (typically within their ASN) when using hosted service, or from the customer's own IP range when self-hosted. The bot supports conditional GET via If-Modified-Since and ETag headers to reduce bandwidth usage, and can follow links automatically. Newer versions can be configured to crawl JavaScript-rendered content using headless browsing. The default request frequency is moderate, but administrators may increase it significantly for large sites, making rate-limiting advisable.
Thunderstone explicitly states that webinator-wbi fully honors all robots.txt directives, including Disallow, Allow, and Crawl-delay. It also respects the Robots Exclusion Protocol standard and the X-Robots-Tag header for non-HTML content. Compliance is hard-coded into the Webinator software, and site operators can verify this through the official manual.
The primary User-Agent string is webinator-wbi, often with a version suffix such as webinator-wbi/4.0 or webinator-wbi/5.0. Common formats include Mozilla/5.0 (compatible; webinator-wbi; +http://www.thunderstone.com/texis/webinator/). Some installations include the administrator's email in a From header. The bot supports Accept-Encoding (gzip) and Connection: keep-alive, but does not send unusual or spoofed headers. The presence of the token "webinator-wbi" is definitive for identification.
Collected content is used solely to populate Webinator's search index, providing full-text search, faceted navigation, and relevance ranking for websites and intranets. No data is repurposed for AI training, machine learning, or sold to third parties. Indexed data remains under the control of the organization running Webinator; analytics are limited to search performance metrics internal to the platform.
Rate limiting of webinator-wbi is recommended because its crawl speed is entirely configurable by the operator, potentially causing high request volumes that impact server performance. Threshold-based blocking ensures fair resource allocation while still permitting legitimate indexing activity, aligning with standard webmaster best practices for non-malicious crawlers.
Similar Threats
🛡️
Boteraser automatically detects and blocks unwanted bots — protecting your site from scrapers, DDoS bursts, and credential stuffing attacks without slowing down real visitors.
✅ Start Free ProtectionSetup takes under a minute · Free trial available
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.
Stay up to date with the latest from Boteraser.
We use cookies to improve your experience on our site. By using our site, you consent to cookies.
Manage your cookie preferences below:
Essential cookies enable basic functions and are necessary for the proper function of the website.
CloudFlare provides web performance and security solutions, enhancing site speed and protecting against threats.
Service URL: developers.cloudflare.com (opens in a new window)
These cookies are needed for adding comments on this website.
These cookies are used for managing login functionality on this website.
Statistics cookies collect information anonymously. This information helps us understand how visitors use our website.
Google Analytics is a powerful tool that tracks and analyzes website traffic for informed marketing decisions.
Service URL: policies.google.com (opens in a new window)
You can find more information in our Cookie Policy and Privacy Policy.