zotag search
Search Engine User-Agent:zotag-search
🤖 Overview
Zotag Search is a web crawler operated by the Corporation for Digital Scholarship, the non‑profit that develops the open‑source reference manager Zotero. Its primary purpose is to automatically discover and index publicly available scholarly content, including journal articles, preprints, and academic web pages, in order to populate Zotero’s automatic tagging and metadata inference system. The data it collects feeds into the Zotag feature, which uses machine‑learning models to suggest subject‑specific tags for users’ reference libraries, thereby enhancing discoverability and organization within Zotero.
🌐 Technical Behavior
Zotag Search employs a controlled crawl strategy that prioritizes publisher websites, institutional repositories, and preprint servers (e.g., arXiv, PubMed Central). Requests are made using standard HTTP/1.1 with a default interval of several seconds between successive fetches to avoid overwhelming servers, as documented in Zotero’s public server configuration. The crawler identifies itself via the User-Agent header as Mozilla/5.0 (compatible; Zotero/2.0; +https://www.zotero.org/) and also uses a dedicated ZotagSearch/1.0 token when specifically collecting data for tag generation. IP addresses originate from the Corporation’s own infrastructure, typically within the 128.118.0.0/16 range (based on ASN AS14325). It follows Link headers and sitemap XML files to discover new content, and it respects Retry-After headers sent by servers under load.
📋 robots.txt Compliance
Official Zotero documentation states that the crawler fully honors robots.txt directives, including Disallow and Crawl-delay instructions. The Zotero development team explicitly advises server administrators to use Disallow: / to exclude all content from Zotero indexing. In practice, the crawler has been observed to adhere strictly to these rules, and any reports of non‑compliance are treated as bugs.
🔍 Detection Indicators
Primary detection is via the User‑Agent string ZotagSearch/1.0 or the broader Zotero/2.0 agent. The crawler also includes a standard From header containing the email address zotero‑[email protected]. Behavioral fingerprints include fetching both the HTML page and any linked PDF or supplementary data files within the same session, and making requests primarily to paths containing /publisher/, /article/, or /doi/. The Accept header is typically text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8, and Accept-Language is set to en‑US,en;q=0.5.
📊 Data Usage
Collected content is used solely within the Zotero ecosystem to extract textual features for automatic tag generation. The tags are suggested to users when they add references to their personal libraries; no raw content is stored beyond what is required for model inference. According to Zotero’s privacy policy, the crawler does not collect personal data or login‑protected content. Tagging models are trained on publicly available web corpora, not on individual user libraries.
⚙️ Rate Limiting Policy
Although Zotag Search is legitimate and well‑behaved, it is rate‑limited by many academic publishers because its sustained crawl volume can still place moderate load on content delivery systems. A threshold‑based block (e.g., more than 10 requests per second from the same IP) is applied to protect server stability while still allowing the crawler’s valuable metadata‑gathering function to operate efficiently.
Similar Threats
⚠️
Your Site May Be Hemorrhaging Revenue to Bots
Unwanted bots inflate your analytics, drain server resources, and slow down real users. Check if your site is affected — completely free.
Check My Site for FreeFree to start · Cancel anytime
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.