libcrawl
libcrawl is a non‑profit web crawler operated by libcrawl.org, established in 2019 to collect open web content for academic AI training and linguistic research. Its dataset feeds into large‑scale language model projects and is available under a Creative Commons license.
libcrawl uses a distributed Scrapy architecture, sending 1–5 requests per second per IP with a default 2‑second crawl delay. It identifies via User‑Agent libcrawl/1.0 (+https://libcrawl.org/bot) and uses HTTP/1.1 with ETag caching. The crawler operates from IPv4 addresses in ASN 12345 and IPv6 prefix 2001:db8::/32. It respects noindex metas and nofollow links, and does not execute JavaScript.
libcrawl fully respects robots.txt directives. It caches robots.txt for 24 hours and never visits Disallowed paths, as documented at libcrawl.org/docs/robots. No violations have been reported.
The primary User‑Agent string is libcrawl/1.0 or libcrawl/2.0. Additional headers include X-Crawl-ID and From: [email protected]. Its IPs reverse‑resolve to crawler-*.libcrawl.org. The bot may be identified by its steady request rate and lack of randomness in intervals.
Text and metadata extracted from crawled pages are used to train transformer models and benchmark information retrieval systems. libcrawl releases annual WARC snapshots on Zenodo and BitTorrent for non‑commercial research use.
Despite its good behavior, libcrawl’s sustained crawl rate can strain small sites. Rate‑limiting at 10 requests/second per IP or when 404 errors exceed 5% is recommended to protect server resources while allowing the crawler to complete its mission.
⚠️
Unwanted bots inflate your analytics, drain server resources, and slow down real users. Check if your site is affected — completely free.
Check My Site for FreeFree to start · Cancel anytime
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.
Stay up to date with the latest from Boteraser.
We use cookies to improve your experience on our site. By using our site, you consent to cookies.
Manage your cookie preferences below:
Essential cookies enable basic functions and are necessary for the proper function of the website.
CloudFlare provides web performance and security solutions, enhancing site speed and protecting against threats.
Service URL: developers.cloudflare.com (opens in a new window)
These cookies are needed for adding comments on this website.
These cookies are used for managing login functionality on this website.
Statistics cookies collect information anonymously. This information helps us understand how visitors use our website.
Google Analytics is a powerful tool that tracks and analyzes website traffic for informed marketing decisions.
Service URL: policies.google.com (opens in a new window)
You can find more information in our Cookie Policy and Privacy Policy.