Anarchy99

Bot User-Agent: anarchy99

🤖 Overview

Anarchy99 is a web crawler operated by Anarchy Labs (anarchylabs.com), a data services firm specializing in large‑scale web scraping for AI training datasets and competitive intelligence. First documented in early 2023, its purpose is to collect publicly available web content to feed the "Anarchy Datasets" platform used by enterprise clients for text classification and market analysis.

🌐 Technical Behavior

The crawler employs a distributed architecture spanning multiple cloud providers, including AWS (us‑east‑1, eu‑west‑1), DigitalOcean (nyc3, sfo2), and Linode. It supports HTTP/1.1 and HTTPS, but does not use HTTP/2 or QUIC. Typical request rates average 5–10 requests per second per IP, with occasional bursts up to 30 requests per second depending on crawl objectives. Anarchy99 follows robots.txt rules recursively but avoids binary content (PDFs, images, videos) unless explicitly allowed. It respects If‑Modified‑Since and Etag headers to reduce redundant downloads, and uses a random delay between page navigations to appear human‑like.

📋 robots.txt Compliance

According to Anarchy Labs’ official documentation at docs.anarchylabs.com/crawlers, the bot fully honors Disallow directives and will not crawl forbidden paths. However, the developer notes explicitly state that Crawl‑Delay is not consistently enforced; the bot’s own internal rate limiter overrides external directives. This policy is publicly documented but may lead to higher than expected request rates on sites with strict crawl‑delay settings.

🔍 Detection Indicators

The primary User‑Agent string is "Anarchy99/1.0" with variants like "Anarchy99‑Bot/1.0" and "Anarchy99‑Preview/1.0" for mobile devices. Additional fingerprints include a custom HTTP header X‑Anarchy‑Token with a hash value derived from the crawl job ID. Reverse DNS lookups on its IPs often resolve to hostnames such as crawler‑*.anarchylabs.net. The bot does not use standard robot‑check headers like X‑Forwarded‑For; instead it sends direct source IPs.

📊 Data Usage

Collected content is used to train proprietary AI models for sentiment analysis, entity extraction, and topic modeling. The derived datasets are sold as subscription‑based APIs or raw training corpora to enterprise customers in finance, e‑commerce, and media sectors. Anarchy Labs explicitly states in its privacy policy that raw page text is not resold; only aggregated insights and model weights are commercialized.

⚙️ Rate Limiting Policy

Rate limiting is recommended because Anarchy99 can sustain high request volumes over long crawl sessions, potentially overwhelming smaller web servers. A threshold of 100 requests per minute per IP, or 10 requests per second with a burst limit of 30, provides a safety margin while allowing legitimate data collection. This aligns with best practices for any aggressive but legitimate crawler.

53% of Web Traffic Is Bots in 2026

— Imperva Bad Bot Report 2026

How much of your traffic is automated? Get your personal bot traffic report and see exactly what's hitting your server — completely free.

📊 Get My Bot Report

Sign up in seconds  ·  No card required

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.