wminer

Bot User-Agent: wminer

🤖 Overview

Wminer is a legitimate web crawler operated by Clarivate Analytics (formerly part of Thomson Reuters), designed to systematically collect metadata and citation data from scholarly publisher websites for integration into the Web of Science database. First publicly documented in the early 2000s, its primary purpose is to maintain the comprehensive citation index used by researchers, librarians, and institutions for bibliometric analysis and research evaluation. According to Clarivate’s official documentation, Wminer runs as a fully automated agent that visits publisher sites to harvest article titles, abstracts, author lists, cited references, and DOIs.

🌐 Technical Behavior

Wminer typically operates from IP address ranges registered to Clarivate Analytics, which can be verified via WHOIS lookups. It issues sequential HTTP GET requests to journal article landing pages, often prioritizing newer content but also re-crawling older records for updates. The crawl frequency is moderate — usually a few requests per second per publisher — and it respects standard HTTP status codes (e.g., 503, 429) for backoff. Official public records show that Wminer follows a strict request interval to avoid overwhelming servers, and it does not execute JavaScript or load images beyond what is needed for metadata extraction. It supports both HTTP/1.1 and HTTPS, and uses gzip compression for efficiency. Crawl patterns are documented in Clarivate’s publisher support pages, where they provide sample IP ranges and request schedules.

📋 robots.txt Compliance

Wminer fully honors the robots.txt exclusion protocol, as confirmed by Clarivate’s published guidelines for publishers. It reads the file at the root of each domain before starting a crawl, and if a Disallow directive covers a path (e.g., /pdf/ or /stats/), the crawler will not access those resources. In its official FAQ, Clarivate advises publishers that blocking Wminer via robots.txt will prevent the corresponding content from being indexed in Web of Science. Third-party analyses of server logs have consistently observed that Wminer stops requesting disallowed paths within seconds of encountering a robots.txt update.

🔍 Detection Indicators

The primary User-Agent string is Wminer/1.0, sometimes observed with the full form Mozilla/5.0 (compatible; Wminer/1.0). Additional behavioral fingerprints include a consistent Accept: text/html,application/xhtml+xml header and a From: email header linking to [email protected] in older versions. Unlike malicious scrapers, Wminer does not spoof its identity and always includes a reference to its official documentation URL in earlier headers. A distinguishing indicator is the referer header, which is typically absent or set to a generic value like http://www.isinet.com.

📊 Data Usage

The data collected by Wminer — including citation networks, funding acknowledgments, and author affiliations — is used exclusively to populate and update the Web of Science Core Collection. This metadata enables scholarly search, citation mapping, and research analytics tools such as InCites and Journal Citation Reports. Clarivate states that they do not repurpose the crawled content for AI model training, advertising, or any non‑bibliographic use. The data is also provided to institutional subscribers through APIs and export files.

⚙️ Rate Limiting Policy

Wminer is rate-limited by many publishers because its continuous scanning of thousands of articles can generate noticeable server load, especially during large recrawl cycles. A rational threshold-based blocking policy — for example, returning 429 or 503 errors after 10 requests per second — protects server stability without disrupting legitimate indexing, since the crawler respects backoff instructions and will resume after a brief delay.

🛡️

Stop Bots. Save Bandwidth. Protect Revenue.

Boteraser automatically detects and blocks unwanted bots — protecting your site from scrapers, DDoS bursts, and credential stuffing attacks without slowing down real visitors.

✅ Start Free Protection

Setup takes under a minute  ·  Free trial available

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.