eccp-1-0
ECCP/1.0 is a web crawler operated by Common Crawl, a nonprofit that provides a free, open repository of web crawl data. Deployed in late 2022 under the Efficient Crawling and Caching Protocol (ECCP), its purpose is to collect public web pages for the Common Crawl dataset, used by researchers and companies for AI training and NLP. The bot is part of an experimental initiative to reduce bandwidth impact while maximizing coverage.
ECCP/1.0 performs HTTP GET requests at a rate of about 5 requests per second per IP, with bursts capped at 50 per minute. It uses IPv4 range 54.201.0.0/16 and IPv6 range 2600:1f18::/32 (per Common Crawl’s IP list at commoncrawl.org/crawler-ips). It supports conditional GET via ETag and Last-Modified, does not fetch JavaScript, and limits itself to HTML and plain text. The bot includes a contact email in headers: [email protected].
According to Common Crawl’s policy on github.com/commoncrawl/cc-crawl-statistics, ECCP/1.0 fully honors robots.txt directives, including Disallow and Crawl-Delay. There are no documented cases of it ignoring exclusions.
The primary User-Agent string is ECCP/1.0. Behavioral fingerprints include a 200ms request interval, preference for .html files, and absence of Accept-Encoding. Requests originate from ASN 14618 (AWS EC2) and include a From: header with [email protected].
Collected data feeds the Common Crawl open dataset, used for training LLMs (e.g., GPT, BERT), building search indexes, and academic research. All data is publicly accessible without cost.
Though legitimate, ECCP/1.0’s volume can strain small servers. Rate limiting with a threshold of 100 requests per minute per IP is a common policy to prevent overload while allowing cooperative crawling.
Similar Threats
🛡️
Boteraser automatically detects and blocks unwanted bots — protecting your site from scrapers, DDoS bursts, and credential stuffing attacks without slowing down real visitors.
✅ Start Free ProtectionSetup takes under a minute · Free trial available
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.
Stay up to date with the latest from Boteraser.
We use cookies to improve your experience on our site. By using our site, you consent to cookies.
Manage your cookie preferences below:
Essential cookies enable basic functions and are necessary for the proper function of the website.
CloudFlare provides web performance and security solutions, enhancing site speed and protecting against threats.
Service URL: developers.cloudflare.com (opens in a new window)
These cookies are needed for adding comments on this website.
These cookies are used for managing login functionality on this website.
Statistics cookies collect information anonymously. This information helps us understand how visitors use our website.
Google Analytics is a powerful tool that tracks and analyzes website traffic for informed marketing decisions.
Service URL: policies.google.com (opens in a new window)
You can find more information in our Cookie Policy and Privacy Policy.