caddbot

Bot User-Agent: caddbot

🤖 Overview

Caddbot is a legitimate web crawler operated by Caddi Inc., a company specializing in AI-powered data extraction and knowledge graph construction. First publicly documented in 2022, the bot’s primary purpose is to collect publicly accessible web content to train and improve Caddi’s proprietary large language models and structured data extraction pipelines. According to Caddi’s official bot information page at https://caddi.com/bot, the crawler operates under a clearly defined policy and is used exclusively for internal model development, not for public search indexing or advertising.

🌐 Technical Behavior

Caddbot performs recursive crawling by following HTML anchor links and sitemap.xml files, with a default crawl frequency of roughly one request every 2–3 seconds per domain, though this rate may increase for larger sites. The bot primarily uses HTTP/1.1 and HTTPS protocols and identifies itself via the User-Agent header. Official documentation notes that Caddbot’s IP addresses belong to the 104.28.0.0/16 and 162.159.0.0/16 ranges (Cloudflare-owned IPs) when operating through Caddi’s proxy infrastructure. The bot respects robots.txt rules before each crawl and does not intentionally bypass CAPTCHA or IP rate limits.

📋 robots.txt Compliance

According to Caddi’s published policy, Caddbot strictly honors Disallow directives found in robots.txt files. The official documentation explicitly states: “We respect robots.txt rules and will not crawl any page that is disallowed.” This compliance has been verified by multiple third-party crawl audits referenced in webmaster forums. Caddbot also reads Crawl-Delay directives and adjusts its request interval accordingly.

🔍 Detection Indicators

The primary identifying User-Agent string is Mozilla/5.0 (compatible; Caddbot/1.0; +https://caddi.com/bot). Some variants omit the Mozilla prefix and use simply Caddbot/1.0. The bot also sends a From header with the email [email protected] for contact purposes. Behavioral fingerprinting reveals a consistent pattern of only crawling text/html and application/rss+xml content types, ignoring images, scripts, and CSS files.

📊 Data Usage

All data collected by Caddbot is used exclusively to train Caddi’s AI models for structured data extraction, entity recognition, and knowledge graph building. The company’s privacy policy states that crawled content is stored temporarily, processed for feature extraction, and then discarded—not stored permanently or shared with third parties. Caddbot does not contribute to any public search index or advertising network.

⚙️ Rate Limiting Policy

Rate limiting is recommended for Caddbot only if its default crawl rate (roughly 20–30 requests per minute) exceeds a site’s capacity or interferes with normal traffic. The rationale for threshold-based blocking is to protect server resources while still allowing legitimate AI training data collection, as outlined in Caddi’s own rate-limiting guidelines published at https://caddi.com/bot#rate-limits.

Free Bot Analysis

Is Your Site Under Bot Attack Right Now?

Find out exactly how much of your traffic is automated — and which bots are draining your bandwidth and skewing your analytics.

Run Free Bot Scan →

No credit card required  ·  Results in minutes

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.