special_archiver

Archiver User-Agent: special-archiver

🤖 Overview

The special_archiver is a web crawling agent operated by the Internet Archive, a non‑profit digital library founded in 1996. Its primary purpose is to systematically capture snapshots of public web pages for inclusion in the Wayback Machine and other archival projects such as Archive‑It. This bot is distinct from the general‑purpose ia_archiver and is deployed specifically for high‑priority or time‑sensitive archival requests originating from partners and researchers.

🌐 Technical Behavior

The special_archiver crawls using a combination of HTTP/1.1 and HTTP/2 protocols, sending requests at an average rate of 2‑3 requests per second per domain to minimize disruption. Its IP addresses belong to the ASN 7941 (Internet Archive) and are drawn from the ranges 207.241.224.0/20 and 64.147.96.0/19, though some requests originate from Cloudflare proxied IPs when the live site uses CDN services. The bot respects HTTP ETags and Last‑Modified headers to avoid re‑fetching unchanged content, and it can follow both HTML anchor links and sitemap indexes defined in robots.txt or sitemap.xml files. Crawl depth is typically limited to 3 levels for regular archives, but special directives can extend it.

📋 robots.txt Compliance

The Internet Archive has publicly stated that special_archiver fully respects Robots Exclusion Protocol directives found in /robots.txt. Official documentation at archive.org/about/exclude confirms that the bot reads Disallow rules and also honors Allow overrides. There is no evidence of the bot ignoring crawl‑delay or user‑agent‑specific instructions.

🔍 Detection Indicators

The primary User‑Agent string reported is Mozilla/5.0 (compatible; special_archiver/2.0; +https://archive.org/details/archive‑it). In some deployments, it appends a project identifier such as (archive‑it‑partner‑123). The bot sends a From header containing the email address [email protected] and a Via header indicating the proxy gateway. Log entries often show repeated requests for robots.txt, sitemap.xml, and favicon.ico before any page fetch.

📊 Data Usage

All content captured by special_archiver is stored as WARC (Web ARChive) files and made publicly accessible through the Wayback Machine at web.archive.org. The data is used for long‑term preservation, academic research, cultural heritage documentation, and legal discovery. No AI training or commercial analytics are performed on the collected pages; the sole purpose is archival access.

⚙️ Rate Limiting Policy

Because the bot is legitimate but can generate high volumes of requests during bulk archival missions, administrators should apply threshold‑based rate limiting (e.g., 10 requests per second per IP) to protect dynamic resources while still allowing the bot to fulfill its preservation mission. Rate limiting is not a block but a safeguard against accidental load spikes.

🛡️

Stop Bots. Save Bandwidth. Protect Revenue.

Boteraser automatically detects and blocks unwanted bots — protecting your site from scrapers, DDoS bursts, and credential stuffing attacks without slowing down real visitors.

✅ Start Free Protection

Setup takes under a minute  ·  Free trial available

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.