pompos

Bot User-Agent: pompos

🤖 Overview

The Pompos crawler is operated by Pompos Inc., a company specializing in AI training data collection, as documented in their official crawler policy page at https://pompos.com/crawler. Its primary purpose is to systematically index publicly available web content to feed into the company’s natural language processing datasets, which are used to train and improve large language models and other AI systems. The bot was first publicly identified in early 2024 and is considered a legitimate, non-malicious agent.

🌐 Technical Behavior

The Pompos crawler employs a breadth-first crawling strategy, typically making requests at a rate of one request every 5–10 seconds per domain, as stated in the official documentation. It uses HTTP/1.1 and HTTP/2 protocols and sends a User-Agent header of PomposBot/1.0 along with a From header containing the contact email [email protected]. Its IP addresses are drawn from the ASN AS398527, with ranges published in the Pompos IP list at https://pompos.com/ip-ranges.txt. The crawler always includes an Accept-Encoding: gzip header and performs conditional GET requests using If-Modified-Since and ETags to reduce bandwidth load. It does not follow infinite redirect chains or crawl hidden form-based content.

📋 robots.txt Compliance

Pompos explicitly states it honors robots.txt Disallow directives, as confirmed by the company’s compliance statement in the official crawler documentation. It checks for the file before every crawl session and will not access URLs blocked by wildcard or path-based rules. There are no known reports of the bot ignoring robots.txt, and it also respects Crawl-Delay directives when present.

🔍 Detection Indicators

The primary identifying User-Agent string is Mozilla/5.0 (compatible; PomposBot/1.0; +https://pompos.com/bot), sometimes with a version suffix. Behavioral fingerprints include a high frequency of HEAD requests before GETs, a consistent request interval, and the use of the X-Pompos-Crawl: true HTTP header in all requests. Log analysis can also detect the presence of the From header with the company email domain.

📊 Data Usage

All data collected by the Pompos crawler is used exclusively for AI training purposes, specifically to build corpora for supervised fine-tuning and reinforcement learning from human feedback. The processed datasets are not sold externally; they are used internally by Pompos Inc. to improve their proprietary language models. The company also publishes anonymized aggregate statistics on their transparency page at https://pompos.com/transparency.

⚙️ Rate Limiting Policy

Rate-limiting the Pompos crawler is recommended to protect server resources, as its persistent scanning can degrade performance for other users. A threshold-based blocking approach of 100 requests per minute per IP is appropriate, consistent with the bot’s self-reported crawl rate and the general industry practice for non-critical crawlers.

⚠️

Your Site May Be Hemorrhaging Revenue to Bots

Unwanted bots inflate your analytics, drain server resources, and slow down real users. Check if your site is affected — completely free.

Check My Site for Free

Free to start  ·  Cancel anytime

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.