henkbot
HenkBot is a web crawler operated by Henk van der Heijden, an independent Dutch software engineer and researcher, first publicly documented in January 2022 on his academic website. Its purpose is to collect publicly accessible web content to analyze web graph connectivity, link distributions, and the evolution of page rank. The data feeds into a non-commercial research database called the Web Graph Archive, used solely for peer-reviewed academic studies.
HenkBot employs a configurable crawl delay, defaulting to 2 seconds between requests, and respects a daily crawl budget of 1000 pages per domain. It operates single‑threaded over IPv4 only, sending standard HTTP/1.1 GET requests from a small pool of IP addresses assigned to a Dutch residential ISP (AS20495). The bot fetches only text/html and application/xhtml+xml content, ignoring images, CSS, and JavaScript. It sets an X-Request-ID header for request correlation and limits its crawl window to 08:00–20:00 UTC. Redirects to non‑HTTP protocols are not followed.
According to the official GitHub repository (github.com/henk/henkbot), the code parses robots.txt before every crawl session and strictly enforces all Disallow directives. It also honors the Crawl-Delay directive if present, with a minimum of 10 seconds as per the developer’s default configuration.
The primary User‑Agent string is HenkBot/1.0 (compatible; Henk Bot; +https://henkbot.example.com), occasionally seen as HenkBot/2.0. Behavioral fingerprints include a fixed Accept-Language header of nl‑NL,en, no cookie acceptance, and request bursts of 1–2 per minute during active crawl windows.
Collected data is used to construct large‑scale directed graph models, study linking patterns and hub‑authority relationships, and monitor structural changes over time. Results are published through the Web Graph Archive and in academic papers available on arXiv. No raw data is sold or shared with third parties.
Although HenkBot is designed to be polite, early versions occasionally experienced misconfigurations that caused rapid bursts across multiple domains. Therefore, rate limiting at 10 requests per minute is recommended to prevent unintended load, while still allowing legitimate academic crawling.
Similar Threats
Free Bot Analysis
Find out exactly how much of your traffic is automated — and which bots are draining your bandwidth and skewing your analytics.
Run Free Bot Scan →No credit card required · Results in minutes
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.
Stay up to date with the latest from Boteraser.
We use cookies to improve your experience on our site. By using our site, you consent to cookies.
Manage your cookie preferences below:
Essential cookies enable basic functions and are necessary for the proper function of the website.
CloudFlare provides web performance and security solutions, enhancing site speed and protecting against threats.
Service URL: developers.cloudflare.com (opens in a new window)
These cookies are needed for adding comments on this website.
These cookies are used for managing login functionality on this website.
Statistics cookies collect information anonymously. This information helps us understand how visitors use our website.
Google Analytics is a powerful tool that tracks and analyzes website traffic for informed marketing decisions.
Service URL: policies.google.com (opens in a new window)
You can find more information in our Cookie Policy and Privacy Policy.