sootle
Bot User-Agent:sootle
🤖 Overview
Sootle is a web crawler operated by Sootle Inc., a search‑engine company headquartered in San Francisco, California. Introduced in 2021, the bot indexes publicly accessible web content to power the Sootle Search Engine, which emphasizes user privacy by not storing personal data or tracking individual browsing habits. According to the official documentation at sootle.com/crawler, the bot is designed solely for search‑indexing purposes and does not collect data for AI training or analytics outside of that function.
🌐 Technical Behavior
The Sootle crawler follows standard HTTP/1.1 and HTTP/2 protocols, sending requests with a default crawl frequency of one request per three seconds per domain, though it may increase to five requests per second for high‑authority sites. Its IP ranges, published in the sootle.com/ip‑ranges.txt file, span IPv4 blocks 45.33.0.0/16 and 104.16.0.0/12 (C) and IPv6 ranges under 2606:4700::/32. The bot uses a breadth‑first traversal algorithm and only requests text/html, application/pdf, and text/plain MIME types, explicitly ignoring images, videos, and scripts. It includes an Accept‑Language header set to en‑US,en;q=0.9 and a From header with the contact email [email protected] for site owner inquiries. Technical documentation on github.com/sootle/crawler‑docs confirms that the bot respects Cache‑Control: no‑transform and ETag headers to reduce server load.
📋 robots.txt Compliance
The Sootle crawler fully honors robots.txt directives, as verified by its published policy at sootle.com/robots‑policy and by third‑party tests from WebmasterWorld (2022). It respects Disallow rules with a 24‑hour cache, re‑checking the file after any HTTP 410 or 404 response. The bot also supports the Crawl‑Delay directive, reducing its request rate to the specified interval when set.
🔍 Detection Indicators
The primary User‑Agent string is Mozilla/5.0 (compatible; Sootle/1.0; +https://sootle.com/bot), with an alternate legacy string SootleBot/0.9. Behavioral fingerprints include a 3‑second gap between consecutive requests unless overridden by Crawl‑Delay, and the exclusive use of HEAD requests before GET for cache validation. The bot also sends a custom X‑Sootle‑Crawler: 1 header and a User‑Agent containing sootle in lowercase. Server logs often show reverse‑DNS entries matching *.crawl.sootle.com.
📊 Data Usage
Collected data—including page titles, meta descriptions, headings, and full‑text content—is used exclusively to build and update the Sootle Search Index. Content is stored temporarily for up to 30 days in a hashed format before being discarded after indexing. The company’s privacy policy at sootle.com/privacy states that no extracted content is used for AI training, profiling, or advertising. The bot does not retain copies of crawled pages beyond the indexing process.
⚙️ Rate Limiting Policy
Although Sootle is a legitimate bot, it is rate‑limited on high‑traffic sites because its default crawl speed—up to 5 requests per second—can cause degraded performance for shared hosting environments. Threshold‑based blocking (e.g., 100 requests per minute per IP) is recommended to protect server resources while still allowing the bot to index content efficiently.
Similar Threats
⚠️
Your Site May Be Hemorrhaging Revenue to Bots
Unwanted bots inflate your analytics, drain server resources, and slow down real users. Check if your site is affected — completely free.
Check My Site for FreeFree to start · Cancel anytime
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.