arquivo-pt
The Arquivo.pt crawler is operated by the Foundation for Science and Technology (FCT) of Portugal, as part of the national web archive project. Its primary purpose is to systematically collect and preserve publicly accessible web content from the .pt domain and Portuguese‑related sites, maintaining a historical repository accessible at https://arquivo.pt.
The crawler uses a custom HTTP client that supports both HTTP/1.1 and HTTPS, and follows the Robots Exclusion Protocol and crawl‑delay directives. Crawl frequency is generally moderate and obeys any specified delay; during large archiving campaigns the request rate may increase. IP addresses originate from FCT’s allocated blocks, typically within the 193.136.0.0/16 and 194.65.0.0/16 ranges. The crawler identifies itself with the User‑Agent string Mozilla/5.0 (compatible; Arquivo.pt/1.0; +https://arquivo.pt/crawler.html) and also as ArquivoWebCrawler/1.0.
Official documentation from https://arquivo.pt/crawler.html states that the crawler fully respects Disallow directives and the Crawl‑Delay rule. It also supports the extended Archive‑It crawl‑delay syntax. No evidence of ignoring robots.txt has been documented; site owners can block the bot entirely via `User‑agent: Arquivo.pt`.
The primary User‑Agent string is Mozilla/5.0 (compatible; Arquivo.pt/1.0; +https://arquivo.pt/crawler.html); a secondary pattern is Mozilla/5.0 (compatible; ArquivoWebCrawler/1.0). Behaviourally, requests come from Portuguese IP ranges and often include an Accept: text/html header. The bot does not send a Referer header and typically uses a GET method for every resource.
All crawled data is stored in the Arquivo.pt repository for long‑term preservation and is made publicly available via the Wayback‑style interface at https://arquivo.pt. The archive supports historical research, cultural heritage preservation, and legal deposit requirements mandated by Portuguese law.
Although Arquivo.pt is a legitimate archiving agent, its campaigns can temporarily generate high request volumes. Rate limiting with threshold‑based blocking is recommended to protect server resources without permanently barring the crawler, which is essential for the archive’s mission.
Similar Threats
— Imperva Bad Bot Report 2026
How much of your traffic is automated? Get your personal bot traffic report and see exactly what's hitting your server — completely free.
📊 Get My Bot ReportSign up in seconds · No card required
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.
Stay up to date with the latest from Boteraser.
We use cookies to improve your experience on our site. By using our site, you consent to cookies.
Manage your cookie preferences below:
Essential cookies enable basic functions and are necessary for the proper function of the website.
CloudFlare provides web performance and security solutions, enhancing site speed and protecting against threats.
Service URL: developers.cloudflare.com (opens in a new window)
These cookies are needed for adding comments on this website.
These cookies are used for managing login functionality on this website.
Statistics cookies collect information anonymously. This information helps us understand how visitors use our website.
Google Analytics is a powerful tool that tracks and analyzes website traffic for informed marketing decisions.
Service URL: policies.google.com (opens in a new window)
You can find more information in our Cookie Policy and Privacy Policy.