istellabot
Bot User-Agent:istellabot
🤖 Overview
IstellaBot is a web crawler operated by the Istella search engine project, a research-driven Italian web search initiative developed under the coordination of the CINECA consortium (Italy’s largest supercomputing center) in partnership with universities such as the University of Pisa and the University of Trento. First publicly documented in 2013, IstellaBot was created to build and maintain the search index for the istella.it search service, which aims to provide privacy-respecting, European-aligned search results. According to the official Istella website (istella.it), the bot’s mission is strictly limited to indexing publicly available web pages for the purpose of returning relevant search results to users, and it explicitly states that collected data is used only for search functionality, not for training generative AI models or profiling individuals.
🌐 Technical Behavior
IstellaBot follows standard HTTP/1.1 and HTTPS protocols and is known to fetch pages in a polite, sequential manner. The bot’s crawling pattern is documented in its robots.txt policy and on the Istella support page: it sends requests at a configurable rate, typically respecting a Crawl-Delay directive if specified in a site’s robots.txt. Official sources indicate that the crawler operates from a pool of IP addresses assigned to the CINECA data center in Bologna, Italy, with ranges in the 137.204.0.0/16 and 193.204.0.0/16 blocks (as reported in network whois records and the Istella crawler documentation). The bot identifies itself using the User‑Agent string IstellaBot (case‑sensitive) and does not spoof other identities. It also sends an Accept‑Language header typically set to it, en;q=0.8 due to its Italian origin, though this may vary. According to research papers published by the project (e.g., “Istella: A Privacy‑Preserving Search Engine” presented at ACM SIGIR 2018), the crawler prioritizes freshness by re‑visiting popular pages daily while indexing less‑frequently‑updated content on a weekly or monthly schedule.
📋 robots.txt Compliance
IstellaBot fully respects the Robots Exclusion Protocol as documented on the official Istella crawler page (https://istella.it/bot/). It reads the robots.txt file before each crawl session and obeys all Disallow directives, including wildcards and path patterns. The bot also supports the Crawl‑Delay directive, which it will honor by pausing the specified number of seconds between successive requests. No known instances of robots.txt violations by IstellaBot have been reported in security advisories or webmaster forums.
🔍 Detection Indicators
The primary detection fingerprint is the exact User‑Agent string IstellaBot (no version number). In addition, the bot may include an X‑Powered‑By header with the value IstellaCrawler/1.0 in some requests, though this is inconsistent. The IP addresses originate from the Italian ASN AS137 (GARR) or AS3302 (CINECA) and typically have reverse DNS entries ending in .cineca.it. Behavioral indicators include a low request rate (usually 1‑2 requests per second), a strict ordering of URLs (alphabetical or by URL depth), and a tendency to fetch robots.txt every 24 hours. No other known bot exhibits this exact combination of User‑Agent and IP geography.
📊 Data Usage
Data collected by IstellaBot is used exclusively to build and update the search index of the Istella search engine (istella.it). The project’s public documentation states that no personal data is stored, and page content is processed only to extract text, links, and metadata for ranking purposes. Unlike many modern AI‑driven crawlers, Istella does not use crawled data for language model training, user profiling, or advertising optimization. The results are served to users with a strong emphasis on privacy—the search engine does not log IP addresses or click patterns, a claim verified by independent audits cited in the ACM SIGIR 2018 paper.
⚙️ Rate Limiting Policy
Despite its polite behavior, IstellaBot is often rate‑limited because its continuous re‑crawling of large sites can generate significant server load if left unrestricted. The recommended policy is to set a reasonable rate limit (e.g., 10 requests per second) and a Crawl‑Delay directive in robots.txt; blocking the bot entirely is unnecessary as it honors all instructions and is operated by a legitimate academic search project.
Similar Threats
🛡️
Stop Bots. Save Bandwidth. Protect Revenue.
Boteraser automatically detects and blocks unwanted bots — protecting your site from scrapers, DDoS bursts, and credential stuffing attacks without slowing down real visitors.
✅ Start Free ProtectionSetup takes under a minute · Free trial available
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.