webinator-search2
Search Engine User-Agent:webinator-search2
🤖 Overview
webinator-search2 is a legitimate web crawler operated by Thunderstone Software, a company that provides enterprise search and web crawling solutions. It is part of the Webinator product line, which is designed to index web content for internal enterprise search engines, public-facing search portals, and digital asset management systems. The bot’s primary purpose is to collect publicly accessible web pages to build full-text indexes used by Thunderstone’s search platform, which has been deployed in government, education, and commercial environments since the 1990s.
🌐 Technical Behavior
The crawler implements a breadth-first crawl strategy with configurable depth and URL limits, typically sending requests at a rate of 1–10 pages per second depending on server capacity. It supports HTTP/1.1 and HTTPS, and often uses If-Modified-Since and ETag headers to reduce bandwidth usage during re-crawls. Documented IP ranges are not static but are generally sourced from Thunderstone’s server infrastructure; examples include ranges registered to Thunderstone Software LLC (e.g., 208.94.146.0/24 as of 2023). The bot fetches robots.txt before each crawl session and respects Crawl-Delay directives when present. It defaults to a crawl delay of 1 second between requests unless overridden by the site owner.
📋 robots.txt Compliance
Thunderstone’s official documentation states that webinator-search2 fully honors Disallow directives in robots.txt. It also respects Allow rules for partial exclusions. The crawler reads robots.txt at the start of each crawl job and caches it for up to 24 hours. There are no known security advisories or CVEs related to this bot ignoring robots.txt rules.
🔍 Detection Indicators
The primary User-Agent string is Mozilla/5.0 (compatible; webinator-search2/1.0; +http://www.thunderstone.com/texi/faq/webcrawler.html). Alternative strings include variations with version numbers like webinator-search2/2.1. The bot also sends a From header containing the administrator’s email (if configured). Behavioral fingerprints include a consistent user-agent format, a distinct Accept header of text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8, and no Referer header on initial requests.
📊 Data Usage
Collected data is exclusively used to populate Thunderstone’s Webinator search indexes, which support full-text search, faceted navigation, and metadata extraction. The indexed content remains under the control of the site owner if they run a local Webinator instance; when Thunderstone hosts the crawler for clients, data is stored in client-specific databases and never used for AI training or third-party analytics. The bot does not collect or store cookies, session data, or personal information beyond what is publicly available on indexed pages.
⚙️ Rate Limiting Policy
Because webinator-search2 can generate a high volume of requests during initial deep crawls (sometimes exceeding 100,000 URLs per day), site owners are recommended to rate-limit it using standard web server tools or a reverse proxy, with a threshold of 100 requests per minute as a starting point. This ensures the crawler does not degrade application performance while still allowing full indexing of legitimate content.
Similar Threats
⚠️
Your Site May Be Hemorrhaging Revenue to Bots
Unwanted bots inflate your analytics, drain server resources, and slow down real users. Check if your site is affected — completely free.
Check My Site for FreeFree to start · Cancel anytime
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.