arquivo-web-crawler
Crawler User-Agent:arquivo-web-crawler
🤖 Overview
The arquivo-web-crawler is a web crawling agent operated by the Foundation for Science and Technology (FCT) of Portugal as part of the Arquivo.pt project, a national web archive that has been preserving Portuguese web content since 1996. Its primary purpose is to systematically discover, download, and archive publicly accessible web pages, with a focus on .pt domain sites and content hosted on Portuguese servers, to provide long-term digital preservation and public access via the Arquivo.pt portal (arquivo.pt). The crawler follows the Nutch-based Heritrix crawler patterns, adapted for high-frequency harvesting of the Portuguese web.
🌐 Technical Behavior
The bot uses a client-server architecture based on the Heritrix 3.x open-source web crawler framework, with custom modifications for incremental crawling and large-scale archiving. It employs a polite crawling strategy with a default request delay of several seconds between consecutive requests to the same host, but during intensive collection cycles (e.g., quarterly snapshots), it may reduce this delay to as low as 2 seconds per host. The crawler originates from IP addresses registered under the Portuguese Academic and Research Network (FCCN), typically within the ranges 193.136.0.0/16 and 194.210.0.0/16, though it also uses a few assigned IPs from outside Portugal for redundancy. It follows HTTP/1.1 GET requests with standard Accept headers and does not use JavaScript rendering; it only fetches plain HTML, CSS, and linked documents (PDFs, images) but skips large binaries over 10 MB unless explicitly configured otherwise. The crawler honors Cache-Control directives and sets the From header to [email protected] for contact purposes.
📋 robots.txt Compliance
According to official documentation on the Arquivo.pt website, the arquivo-web-crawler strictly adheres to the robots.txt exclusion standard. It checks for Disallow directives before every request and pauses crawls on any path explicitly disallowed, including common exclusions like /cgi-bin/ or administrative directories. The project encourages website owners to use User-agent: arquivo-web-crawler in their robots.txt file to control access.
🔍 Detection Indicators
The definitive User-Agent string for this bot is: Mozilla/5.0 (compatible; arquivo-web-crawler/2.0; +https://arquivo.pt/politicas-de-utilizacao) — note that the string includes the word "arquivo" with a 'q' (Portuguese spelling) and a version number prefixing the contact URL. Additional identifying headers include From: [email protected] and X-Crawler: Arquivo.pt. The bot does not spoof other user agents and always includes its own unique identifier.
📊 Data Usage
All collected data is stored in the Arquivo.pt digital repository and made publicly accessible through the web archive interface at arquivo.pt, where users can search and retrieve historical versions of Portuguese websites. The archive is used for academic research, cultural heritage preservation, and legal compliance with Portuguese national archiving laws. The crawled content is not used for AI training or commercial analytics; it serves purely as a public historical record.
⚙️ Rate Limiting Policy
Although the bot is legitimate and polite, it can become aggressive during large-scale crawls that follow re-crawling schedules (e.g., monthly snapshots of all .pt domains). Rate limiting with a short-term threshold is advisable to prevent saturation of server resources while still allowing the archive to perform its preservation mission.
Free Traffic Analysis
What's Actually Crawling Your Website?
Discover which unwanted bots are being blocked on your site, how often they hit, and where they come from — real data from your own traffic, not guesswork.
🔍 Scan My Site FreePowered by JA4 fingerprinting, honeypot traps & behavioral analysis
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.