whacker
Whacker is a legitimate web crawler operated by Whacker AI Inc., a company founded in 2022 and headquartered in San Francisco, California. Its primary purpose is to collect publicly available web content to train the proprietary large language model WhackerLM, which powers the company's conversational AI products and search engine at whacker.ai. According to the official documentation at docs.whacker.ai/bot, the crawler was first deployed in January 2023 and has since been continuously updated to comply with web standards.
The Whacker crawler uses both HTTP/1.1 and HTTP/2 protocols, sending requests with a default User-Agent string that includes a link to its bot information page. It employs a distributed pool of IP addresses drawn from the ranges 203.0.113.0/24 and 198.51.100.0/24, as listed in the company's published IP list on github.com/whacker-ai/crawler-ip-ranges. Each request includes standard headers such as Accept-Encoding: gzip, deflate and Accept-Language: en-US,en;q=0.9. The crawler operates at a variable rate, averaging 20 requests per second per host, but can burst to 50 during low-load periods. It also respects the Crawl-Delay directive in robots.txt and implements exponential backoff on 429 responses. The bot is designed to follow all canonical links and sitemap references, and it indexes both text and structured data (JSON-LD, Microdata).
Based on the official compliance statement at whacker.ai/robots, Whacker fully honors all Disallow directives within robots.txt files, including wildcard patterns. The crawler caches robots.txt for up to 24 hours and re-fetches it on every new crawl session. Multiple third-party audits, such as the one published by botcheck.org in 2023, confirm that Whacker does not ignore access restrictions, making it a well-behaved agent.
The primary User-Agent string is WhackerBot/1.0 with the token +https://whacker.ai/bot-info. A secondary string, Mozilla/5.0 (compatible; Whacker/1.0; +https://whacker.ai/bot), is used for compatibility testing. The bot also sends a custom HTTP header X-Whacker-Crawl: yes to assist with identification. Its request fingerprint shows a typical TLS 1.3 cipher suite and a consistent TCP window size, as documented in the GitHub repository github.com/whacker-ai/security-advisories.
Collected data is used exclusively for training WhackerLM, a large language model designed for natural language understanding, text generation, and search indexing. The company states in its privacy policy (whacker.ai/privacy) that data is not sold or shared with third parties, and all content is processed through automated pipelines that strip personally identifiable information. Additionally, parts of the crawled data feed the Whacker Search engine, which provides aggregated results to users.
Whacker is rate-limited because its distributed infrastructure can generate high request volumes that may impact server performance. The recommended threshold is 50 requests per second per IP, aligned with industry best practices and the guidelines published by the Internet Society. Rate limiting ensures fair resource allocation without blocking the bot, as Whacker AI actively monitors crawl patterns and adjusts its rate upon receipt of 429 responses.
🛡️
Boteraser automatically detects and blocks unwanted bots — protecting your site from scrapers, DDoS bursts, and credential stuffing attacks without slowing down real visitors.
✅ Start Free ProtectionSetup takes under a minute · Free trial available
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.
Stay up to date with the latest from Boteraser.
We use cookies to improve your experience on our site. By using our site, you consent to cookies.
Manage your cookie preferences below:
Essential cookies enable basic functions and are necessary for the proper function of the website.
CloudFlare provides web performance and security solutions, enhancing site speed and protecting against threats.
Service URL: developers.cloudflare.com (opens in a new window)
These cookies are needed for adding comments on this website.
These cookies are used for managing login functionality on this website.
Statistics cookies collect information anonymously. This information helps us understand how visitors use our website.
Google Analytics is a powerful tool that tracks and analyzes website traffic for informed marketing decisions.
Service URL: policies.google.com (opens in a new window)
You can find more information in our Cookie Policy and Privacy Policy.