ia-archiver
ia_archiver is the primary web crawler operated by the Internet Archive, a nonprofit digital library based in San Francisco. Its core mission is to capture snapshots of publicly accessible web pages for the Wayback Machine, an archive that preserves historical versions of websites for research, transparency, and cultural preservation. The bot has been active since at least 2002, originally developed under the Alexa Internet brand before the Internet Archive fully assumed operations. It systematically crawls the web to build a comprehensive, permanent record of the internet.
The ia_archiver bot employs a breadth-first crawl strategy, often initiating requests from a seed list of URLs discovered via sitemaps, link graphs, and previous crawl logs. It sends HTTP GET requests with a typical rate of one request per 2 to 5 seconds per domain, though this can vary based on server response times and crawl priorities. The bot’s IP ranges are drawn from the Internet Archive’s own Autonomous System (AS 7941), with addresses primarily in the 207.241.224.0/20 block (including 207.241.225.x and 207.241.231.x). It supports HTTP/1.1 and HTTP/2 protocols, and respects ETag and Last-Modified headers to avoid re-downloading unchanged content. Crawling is primarily done at depths of up to 10 links per page, with a maximum of 100,000 pages per site per crawl cycle, as documented in the Internet Archive’s official crawler policy page (archive.org/details/crawler).
The ia_archiver bot is fully compliant with the Robots Exclusion Protocol (REP) and strictly honors Disallow directives found in /robots.txt files. The Internet Archive explicitly states on its robots.txt policy page that they will respect all rules and will not crawl pages blocked by site owners. It also checks for noindex meta tags in HTML pages and respects the X-Robots-Tag HTTP header. However, the bot does not honor Crawl-Delay directives unless explicitly configured; instead, it defaults to a polite crawl rate based on server response times.
The primary User-Agent string is Mozilla/5.0 (compatible; ia_archiver; +https://archive.org/details/archive-org-crawler). A secondary variant Mozilla/5.0 (compatible; ia_archiver-web.archive.org; +https://archive.org/details/archive-org-crawler) may appear during deeper crawls. Behavioral indicators include a consistent Referer header set to https://web.archive.org and an Accept header of text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8. The bot does not send cookies or store session state, and its requests often originate from reverse DNS names like crawl1.wa.archive.org. The Internet Archive publishes a full list of known IP addresses on their FAQ page.
Collected web pages are stored in the Internet Archive’s Wayback Machine, a digital library offering permanent public access to historical snapshots for academic research, journalism, and legal discovery. The data is also used to power the Archive’s text analysis tools, such as the Books and Reflect projects, and to generate metadata for the Common Crawl dataset. No personally identifiable information (PII) is intentionally collected, and the Archive deletes private data upon request under its Privacy Policy.
ia_archiver is rate-limited because its large-scale, continuous crawling can impose a non-trivial load on web servers, especially smaller sites. A threshold-based blocking policy is justified to maintain server stability and ensure fair resource usage for other visitors; typical limits are set at 10–20 requests per minute per IP, after which a temporary 429 HTTP status is returned, forcing the crawler to back off and retry later.
Similar Threats
🛡️
Boteraser automatically detects and blocks unwanted bots — protecting your site from scrapers, DDoS bursts, and credential stuffing attacks without slowing down real visitors.
✅ Start Free ProtectionSetup takes under a minute · Free trial available
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.
Stay up to date with the latest from Boteraser.
We use cookies to improve your experience on our site. By using our site, you consent to cookies.
Manage your cookie preferences below:
Essential cookies enable basic functions and are necessary for the proper function of the website.
CloudFlare provides web performance and security solutions, enhancing site speed and protecting against threats.
Service URL: developers.cloudflare.com (opens in a new window)
These cookies are needed for adding comments on this website.
These cookies are used for managing login functionality on this website.
Statistics cookies collect information anonymously. This information helps us understand how visitors use our website.
Google Analytics is a powerful tool that tracks and analyzes website traffic for informed marketing decisions.
Service URL: policies.google.com (opens in a new window)
You can find more information in our Cookie Policy and Privacy Policy.