digitalarchivesbot
digitalarchivesbot is a web crawler operated by the Internet Archive, a non-profit digital library founded in 1996 by Brewster Kahle, based in San Francisco, California. This bot is specifically used for the Internet Archive's Digital Archives program, which focuses on collecting and preserving publicly accessible web pages, government documents, e-books, and other digital materials for long-term archival storage and public access via the Wayback Machine and other services.
The crawler employs the Heritrix web crawling framework, version 3.x, and uses a breadth-first crawl strategy. It respects HTTP cache-control headers and sends conditional GET requests with If-Modified-Since headers to minimize bandwidth. The bot makes requests at a default rate of approximately one request per second per domain, but this can be adjusted based on site performance. It uses a rotating pool of IP addresses within the Internet Archive's owned ranges, including 207.241.224.0/20 and 64.147.112.0/20. It supports HTTP/1.1 and HTTP/2, and includes headers such as Accept-Encoding: gzip, Accept-Language: en-US,*, and a User-Agent string identifying itself. The crawler also obeys HTTP 429 Too Many Requests responses and will back off accordingly.
According to the Internet Archive's official documentation at archive.org/about/faq.php, digitalarchivesbot fully honors all Disallow directives in a site's robots.txt file. It also respects Crawl-delay directives and will not crawl paths that are disallowed. The bot reads robots.txt at the start of each crawl session and caches it for the duration.
The primary User-Agent string is digitalarchivesbot/1.0 (http://www.archive.org/details/digitalarchivesbot). A variant string digitalarchivesbot/2.0 has also been observed. The bot may include a Via header or a From header with [email protected]. Reverse DNS lookups on its IP addresses resolve to archive.org or wayback.archive.org.
Collected data is used exclusively for the Internet Archive's mission of providing Universal Access to All Knowledge. This includes populating the Wayback Machine for historical web page access, the Digital Archives of Government Websites in partnership with the Library of Congress, and other research collections. The data is also made available for text mining, machine learning, and academic research through the Internet Archive's bulk data download services.
Because digitalarchivesbot can be aggressive when crawling large sites with many pages, rate limiting is implemented to prevent server overload. A threshold-based blocking policy — for example, exceeding 20 requests per second from a single IP address — triggers a temporary block to ensure fair use while still allowing the archiving mission to proceed.
Similar Threats
⚠️
Unwanted bots inflate your analytics, drain server resources, and slow down real users. Check if your site is affected — completely free.
Check My Site for FreeFree to start · Cancel anytime
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.
Stay up to date with the latest from Boteraser.
We use cookies to improve your experience on our site. By using our site, you consent to cookies.
Manage your cookie preferences below:
Essential cookies enable basic functions and are necessary for the proper function of the website.
CloudFlare provides web performance and security solutions, enhancing site speed and protecting against threats.
Service URL: developers.cloudflare.com (opens in a new window)
These cookies are needed for adding comments on this website.
These cookies are used for managing login functionality on this website.
Statistics cookies collect information anonymously. This information helps us understand how visitors use our website.
Google Analytics is a powerful tool that tracks and analyzes website traffic for informed marketing decisions.
Service URL: policies.google.com (opens in a new window)
You can find more information in our Cookie Policy and Privacy Policy.