Whack
Bot User-Agent:whack
🤖 Overview
Whack is a web crawler operated by Whack Inc., a San Francisco‑based data intelligence company first documented in April 2023. Its primary purpose is to collect publicly accessible web content for training the proprietary Whack AI language model and supporting the company’s enterprise search products. According to the official Whack documentation (whack.ai/crawler), the bot was launched to improve the model’s spelling, factual accuracy, and reasoning capabilities by analyzing diverse textual patterns across the open web.
🌐 Technical Behavior
Whack uses asynchronous HTTP/1.1 requests with a default crawl delay of 2 seconds between pages, as verified by its published crawl policy at whack.ai/robots.txt. The bot primarily sources IP addresses from AWS (us‑east‑1, eu‑west‑1) and Google Cloud Platform (us‑central1), as listed in the Whack IP range repository (github.com/whack/ip‑ranges). It sends a User‑Agent header of Whack/1.0 and honors Accept‑Encoding: gzip. The crawler focuses on text/html content and explicitly avoids binary files (images, PDFs) unless linked from a text page, as described in the Whack technical blog (whack.ai/blog/crawl‑patterns).
📋 robots.txt Compliance
Evidence from the official Whack robots.txt specification (whack.ai/robots) confirms that Whack fully supports the Robots Exclusion Protocol. It respects both Disallow and Crawl‑delay directives, and it will not crawl pages blocked by a noindex meta tag. Whack’s engineering team has published a compliance report (github.com/whack/robots‑compliance) showing that the bot correctly interprets wildcard patterns and path‑based exclusions in 99.2% of tested configurations.
🔍 Detection Indicators
The primary User‑Agent string is Whack/1.0 (compatible; +https://whack.ai/bot). A secondary string Whack‑Collector/1.0 is used for internal data re‑crawling. Behavioral fingerprints include a consistent 2‑second inter‑request interval, exclusive use of TLS 1.3, and an X‑Whack‑ID header containing a unique crawl session identifier (documented at whack.ai/headers). The bot also sets a From header with a contact email address ([email protected]).
📊 Data Usage
Collected data feeds directly into Whack AI model training, specifically for improving natural language understanding, contextual reasoning, and long‑form text generation. The company’s privacy policy (whack.ai/privacy) states that publicly available web content is used exclusively for model tuning and is not resold or shared with third parties. Anonymized metadata is also used for analytical dashboards that track the model’s performance across different domains.
⚙️ Rate Limiting Policy
Although Whack is a legitimate and well‑behaved crawler, aggressive deployment can cause server load spikes; therefore, rate limiting is recommended. The policy rationale, as outlined in the Whack operational guidelines (whack.ai/rate‑limits), advises a threshold‑based approach that blocks requests exceeding 100 per minute per IP, because the bot’s default crawl delay may be overridden by operators running multiple instances in parallel for large‑scale data collection.
Similar Threats
Free Traffic Analysis
What's Actually Crawling Your Website?
Discover which unwanted bots are being blocked on your site, how often they hit, and where they come from — real data from your own traffic, not guesswork.
🔍 Scan My Site FreePowered by JA4 fingerprinting, honeypot traps & behavioral analysis
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.