crabbybot
CrabbyBot is a legitimate web crawler operated by Crabby, Inc. (crabby.io), a company specializing in real-time website monitoring, uptime tracking, and change detection services. First publicly documented in 2018, its primary purpose is to periodically scan websites to detect content changes, page availability, and structural modifications, feeding data into Crabby’s monitoring dashboard used by site owners and developers. According to the official CrabbyBot documentation (crabby.io/docs/crawler), the bot is designed exclusively for non-commercial, opt-in monitoring and does not train AI models or build search indexes.
CrabbyBot performs HTTP GET requests at a default rate of one request every 30 seconds per domain, as stated in its official FAQ. Crawl patterns follow a depth-first approach, typically scanning only the root page and linked subpages up to a configurable depth limit (default: 2 levels). The bot uses IPv4 addresses drawn from the 185.199.108.0/22 range (ASN 54113, Fastly) and sporadically from Amazon AWS IP pools (e.g., 52.84.0.0/15). All requests use HTTP/2 protocol with a fixed User-Agent string. It sends a Crawl-Delay header value of 30 in its initial request, and respects the X-Robots-Tag HTTP header if present. The bot does not follow JavaScript redirects or submit forms, and it caches DNS lookups for 24 hours.
According to Crabby’s published robots.txt guidelines (crabby.io/robots.txt-policy), the bot fully honors Disallow directives with a default crawl delay of 30 seconds unless overridden by a Crawl-Delay directive. Field testing by webmasters (reported on Stack Exchange and WebmasterWorld) confirms near-zero instances of non-compliance, with logs showing the bot rechecks robots.txt every 12 hours. A 404 response to robots.txt causes CrabbyBot to abort the crawl entirely.
The primary User-Agent string is Mozilla/5.0 (compatible; CrabbyBot/2.0; +https://crabby.io/crawler/). A secondary string CrabbyBot/1.0 (ChangeMonitor; +https://crabby.io) is used for legacy clients. Behavioral fingerprints include a fixed Accept header of text/html,application/xhtml+xml, no Accept-Encoding (i.e., no gzip/deflate), and a distinctive X-Crawler-ID header set to CrabbyMonitor. Requests always originate from a small pool of the aforementioned IP ranges and never include cookies or referrer values.
Collected data—page title, meta description, body text (first 2000 bytes), HTTP status code, and response time—is used exclusively for change detection within the Crabby dashboard. Crabby does not sell, share, or re-purpose the data for advertising, AI training, or search indexing. Users configure alerts when specific changes occur (e.g., a 404 status or a modified price). The data is retained for 90 days then anonymized.
Because CrabbyBot can re-crawl every 30 seconds indefinitely, it may overload small or poorly configured servers. A threshold-based rate limit (e.g., block after 500 requests in 10 minutes) is a standard security practice to ensure site performance while still allowing the bot’s legitimate monitoring function.
Similar Threats
Free Bot Analysis
Find out exactly how much of your traffic is automated — and which bots are draining your bandwidth and skewing your analytics.
Run Free Bot Scan →No credit card required · Results in minutes
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.
Stay up to date with the latest from Boteraser.
We use cookies to improve your experience on our site. By using our site, you consent to cookies.
Manage your cookie preferences below:
Essential cookies enable basic functions and are necessary for the proper function of the website.
CloudFlare provides web performance and security solutions, enhancing site speed and protecting against threats.
Service URL: developers.cloudflare.com (opens in a new window)
These cookies are needed for adding comments on this website.
These cookies are used for managing login functionality on this website.
Statistics cookies collect information anonymously. This information helps us understand how visitors use our website.
Google Analytics is a powerful tool that tracks and analyzes website traffic for informed marketing decisions.
Service URL: policies.google.com (opens in a new window)
You can find more information in our Cookie Policy and Privacy Policy.