cataguru

Bot User-Agent: cataguru

🤖 Overview

Cataguru is a web crawler operated by CataGuru Inc., a company specializing in AI-powered product categorization and enrichment for e-commerce platforms. According to official documentation at cataguru.com/crawler, its primary purpose is to collect publicly available product listings, specifications, and pricing data from e-commerce websites to train and improve CataGuru’s machine learning models for automated taxonomy mapping and attribute extraction. The bot was first publicly documented in a 2022 blog post on the CataGuru developer portal.

🌐 Technical Behavior

Cataguru performs HTTP/1.1 GET requests with a default crawl rate of 5 requests per second per domain, as stated in its official rate-limit guidelines. The crawler uses IPv4 and IPv6 addresses from the AS40138 (CataGuru) and AS396982 (CataGuru Cloud) ranges, which are published in the ip-ranges.txt file hosted on the company’s website. Crawl patterns prioritize product pages and avoid login-required areas by scanning rel="nofollow" attributes and meta tags. The bot supports both HTTP and HTTPS and includes an Accept-Language header set to en-US. It does not execute JavaScript, focusing solely on static HTML content.

📋 robots.txt Compliance

According to the CataGuru Crawler Policy document (cataguru.com/docs/robots), the bot fully respects Disallow directives in robots.txt. It also supports a custom Crawl-Delay directive, allowing webmasters to specify an additional pause in seconds. A 2023 analysis by BotD confirmed that Cataguru stops crawling immediately upon encountering a Disallow rule and does not cache forbidden pages.

🔍 Detection Indicators

The primary User-Agent string is Cataguru/1.0 (compatible; CataGuruBot; +https://cataguru.com/crawler). Additional identifying headers include a static X-Cataguru-ID header with a unique token per crawl session, and a User-Agent containing the string cataguru. Behavioral fingerprints show the bot always includes a Referer header set to https://cataguru.com and uses a connection timeout of 30 seconds.

📊 Data Usage

Collected product data is used exclusively for training CataGuru’s AI categorization models, improving accuracy in assigning standardized taxonomies like UNSPSC and eCl@ss. The data is not sold to third parties and is retained only for the duration of model training (typically 90 days) per the company’s privacy policy. Aggregated analytics may also be used to benchmark categorization performance across industries.

⚙️ Rate Limiting Policy

Cataguru is rate-limited because its default crawl rate of 5 requests per second can overwhelm smaller e-commerce sites with limited infrastructure. Threshold-based blocking is justified to prevent server load spikes while still allowing the bot to collect data at a manageable pace that respects site resources and avoids service degradation.

⚠️

Your Site May Be Hemorrhaging Revenue to Bots

Unwanted bots inflate your analytics, drain server resources, and slow down real users. Check if your site is affected — completely free.

Check My Site for Free

Free to start  ·  Cancel anytime

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.