iltrovatore
iltrovatore is a web crawler operated by the Wikimedia Foundation, first publicly documented in 2015 as part of the Il Trovatore project. Its primary purpose is to index publicly accessible content from Wikimedia projects (e.g., Wikipedia, Commons, Wiktionary, Wikisource) to power the CirrusSearch engine, which provides full-text search across all Wikimedia sites. The bot is named after the Italian opera “Il Trovatore” and is maintained by the Wikimedia Search team. Official documentation on MediaWiki.org states that the crawler is designed to refresh the search index continuously, ensuring search results reflect the latest edits and additions.
The crawler uses standard HTTP/1.1 requests and typically sends requests at a rate of one page every few seconds, though it can burst to higher frequencies under load. It employs conditional GET requests with If-Modified-Since and If-None-Match headers to minimize bandwidth usage. IP addresses originate from the Wikimedia Foundation’s own infrastructure, which is announced via ASN 14907 and includes ranges such as 208.80.154.0/24 and 2620:0:861:2::/64. The crawler follows a breadth‑first link traversal strategy, respecting the order of links on each page. It also handles redirects and canonical URLs properly. The bot does not execute JavaScript and only fetches HTML, CSS, images, and other static assets when required for indexing.
Wikimedia officially documents that iltrovatore fully honors robots.txt directives, including Disallow and Crawl-Delay rules. The bot’s source code, available on Wikimedia’s GitLab repository (gitlab.wikimedia.org/repos/search/iltrovatore), implements a strict robots.txt parser. It also obeys the X‑Robots‑Tag HTTP header and the noindex meta tag. There are no known instances of intentional violations.
The primary User‑Agent string is: Mozilla/5.0 (compatible; iltrovatore/1.0; +https://www.mediawiki.org/wiki/Il_Trovatore). A secondary variant drops the Mozilla prefix: iltrovatore/1.0 (https://www.mediawiki.org/wiki/Il_Trovatore). The bot does not set custom headers beyond standard ones like Accept and Accept-Encoding. It may also identify itself via the User‑Agent: MediaWiki-CirrusSearch when performing internal requests. Behavioral fingerprints include a consistent request interval, a lack of referrer headers, and a preference for text/html content.
Collected data is used exclusively to build and update the search index for CirrusSearch, which runs on top of Elasticsearch. The index supports full‑text search, phrase matching, and relevance ranking across all Wikimedia sub‑projects. No data is sold, shared with third parties, or used for AI training. The bot does not store raw content longer than necessary for indexing; old index segments are regularly purged.
Although iltrovatore is a legitimate, well‑behaved crawler, it can generate heavy traffic when re‑indexing large sites like Wikipedia (over 6 million articles). Rate limiting is implemented at the web server or CDN layer (e.g., using nginx or Varnish) to cap requests per IP per second, typically at 5–10 requests per second, to prevent resource exhaustion while still allowing the bot to complete its indexing within a reasonable timeframe.
🛡️
Boteraser automatically detects and blocks unwanted bots — protecting your site from scrapers, DDoS bursts, and credential stuffing attacks without slowing down real visitors.
✅ Start Free ProtectionSetup takes under a minute · Free trial available
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.
Stay up to date with the latest from Boteraser.
We use cookies to improve your experience on our site. By using our site, you consent to cookies.
Manage your cookie preferences below:
Essential cookies enable basic functions and are necessary for the proper function of the website.
CloudFlare provides web performance and security solutions, enhancing site speed and protecting against threats.
Service URL: developers.cloudflare.com (opens in a new window)
These cookies are needed for adding comments on this website.
These cookies are used for managing login functionality on this website.
Statistics cookies collect information anonymously. This information helps us understand how visitors use our website.
Google Analytics is a powerful tool that tracks and analyzes website traffic for informed marketing decisions.
Service URL: policies.google.com (opens in a new window)
You can find more information in our Cookie Policy and Privacy Policy.