infomine
Infomine is a web crawler operated by Infomine Technologies, designed to index academic and research-oriented content for the Infomine search engine (www.infomine.com). First deployed in the early 2000s, it focuses on collecting publicly accessible scholarly articles, university resources, and government publications to populate its specialized database.
The crawler follows a breadth-first traversal pattern, requesting pages at intervals of 5–10 seconds, but can scale up to multiple concurrent requests when encountering large sites. It primarily uses HTTP/1.1 and respects standard request headers. Its IP addresses are drawn from a pool belonging to Infomine's own ASN (ASxxxxx) and occasionally from cloud providers, though the full range is not publicly documented. The bot identifies via the User-Agent string "Mozilla/5.0 (compatible; Infomine/1.0; +http://www.infomine.com/bot.html)" and does not spoof other agents.
According to the official bot page (http://www.infomine.com/bot.html), Infomine fully honors robots.txt directives, including Disallow and Crawl-delay settings. Evidence from site administrators confirms that the bot halts crawling on restricted paths within one crawl cycle.
The primary indicator is the User-Agent "Infomine/1.0" with the trailing bot URL. Additionally, the bot includes a custom header "From: [email protected]" in some implementations. It does not exhibit unusual request patterns or hidden identifiers.
Collected data feeds the Infomine search index, which specializes in academic and research materials. The index is used for public search and may also be used for internal analytics and trend monitoring in the scholarly domain. No evidence suggests use for AI model training.
Because the crawler can generate significant request volume on large sites, it is rate-limited to prevent server overload. A threshold of 100 requests per minute per IP is commonly applied by site operators, consistent with standard best practices for aggressive legitimate crawlers.
Free Bot Analysis
Find out exactly how much of your traffic is automated — and which bots are draining your bandwidth and skewing your analytics.
Run Free Bot Scan →No credit card required · Results in minutes
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.
Stay up to date with the latest from Boteraser.
We use cookies to improve your experience on our site. By using our site, you consent to cookies.
Manage your cookie preferences below:
Essential cookies enable basic functions and are necessary for the proper function of the website.
CloudFlare provides web performance and security solutions, enhancing site speed and protecting against threats.
Service URL: developers.cloudflare.com (opens in a new window)
These cookies are needed for adding comments on this website.
These cookies are used for managing login functionality on this website.
Statistics cookies collect information anonymously. This information helps us understand how visitors use our website.
Google Analytics is a powerful tool that tracks and analyzes website traffic for informed marketing decisions.
Service URL: policies.google.com (opens in a new window)
You can find more information in our Cookie Policy and Privacy Policy.