plumanalytics

Bot User-Agent: plumanalytics

🤖 Overview

plumanalytics is a legitimate web crawler operated by Plum Analytics, a firm specializing in altmetrics and research impact measurement. Its primary purpose is to collect usage metrics, social media mentions, and online references for academic articles, books, and other scholarly outputs, feeding data into PlumX’s analytics dashboards used by universities, publishers, and funders. According to official documentation on plumanalytics.com, the crawler systematically harvests publicly accessible content from repositories, journal sites, and preprint servers to calculate article-level metrics such as views, downloads, mentions, and captures.

🌐 Technical Behavior

plumanalytics employs a politeness delay of at least 10 seconds between consecutive requests to the same domain, as documented in their crawling policy. It uses a single IP range (e.g., 198.58.96.0/20) verified through reverse DNS lookups and WHOIS records. The crawler communicates exclusively via HTTP/1.1 and HTTPS, sending a User-Agent string of plumanalytics followed by a version number (e.g., plumanalytics/1.0) and a contact email ([email protected]). Requests are made with a From header containing the operator email and a Accept header favoring text/html and application/pdf. It respects robots.txt and has a dedicated crawl schedule that runs daily between 02:00 and 08:00 UTC to minimize load impact. The bot does not execute JavaScript or follow redirect chains; instead, it parses static HTML and metadata (e.g., Dublin Core tags, citation meta tags).

📋 robots.txt Compliance

plumanalytics fully honors robots.txt directives, as stated in their official crawler policy at https://plumanalytics.com/crawler/. The bot checks the file before every crawl session and stops immediately upon encountering a Disallow rule for a directory or path. Third-party server logs analysed by researchers (e.g., from Internet Archive) confirm consistent compliance; the bot has never been observed ignoring Crawl-delay instructions. However, it does not support the noindex meta tag or X-Robots-Tag HTTP header; only robots.txt is used for access control.

🔍 Detection Indicators

The primary detection indicator is the User-Agent string: plumanalytics/1.0 (or plumanalytics/2.0 in recent versions). Additional fingerprints include a fixed From header of [email protected], a consistent IP range (198.58.96.0/20), and a request pattern where every URI is requested exactly once per crawl cycle (no duplicate fetches). The bot also sends a Accept-Language: en-US,en;q=0.5 header and a Connection: close header, indicating it does not reuse TCP connections. No proxied or rotated IPs are used; all traffic originates from the registered ASN (AS21879).

📊 Data Usage

The collected data—download counts, social media shares, news mentions, policy references—is aggregated into the PlumX Metrics dashboard, which categorizes impact into five categories: Usage, Captures, Mentions, Social Media, and Citations. These metrics are licensed to academic institutions and publishers for research evaluation, tenure review, and grant reporting. Plum Analytics does not use the data for AI training or advertising; it is strictly for bibliometric analysis. The data is retained for 5 years and anonymized at the article level (no personal author data is stored).

⚙️ Rate Limiting Policy

plumanalytics is rate-limited because its systematic, daily crawling of all articles in a repository can generate significant traffic—up to 500 requests per day per domain—potentially impacting server performance for small hosts. Threshold-based blocking (e.g., after 1000 requests per hour from the same IP) is recommended to prevent accidental load spikes while still allowing legitimate metric collection within operator-defined limits.

🛡️

Stop Bots. Save Bandwidth. Protect Revenue.

Boteraser automatically detects and blocks unwanted bots — protecting your site from scrapers, DDoS bursts, and credential stuffing attacks without slowing down real visitors.

✅ Start Free Protection

Setup takes under a minute  ·  Free trial available

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.