tausdatabot
Bot User-Agent:tausdatabot
🤖 Overview
tausdatabot is a web crawler operated by Taus Inc. (taus.com), a data platform specializing in high-quality training datasets for artificial intelligence. First publicly documented in early 2023, the bot systematically collects publicly available web content to feed into Taus’s proprietary data curation pipeline, which then supplies vetted, labeled datasets for training large language models (LLMs) and other AI systems. Taus’s official documentation describes the crawler as a legitimate, non-malicious agent that adheres to standard web crawling protocols and is designed to minimize server load.
🌐 Technical Behavior
According to Taus’s published crawling policy (taus.com/crawling-policy), tausdatabot performs HTTP GET requests at a default rate of one request per 10 seconds per domain, though it dynamically adjusts based on server response times and Retry-After headers. The bot uses IPv4 and IPv6 addresses from a Taus-owned ASN (AS399507) and occasionally routes through shared cloud IP ranges from AWS and GCP. Crawl patterns prioritize text-heavy pages — HTML, JSON-LD, and RSS feeds — and the bot automatically retries failed requests with exponential backoff. Logs indicate tausdatabot sends a X-Taus-Crawler: true header in all requests to aid in identification.
📋 robots.txt Compliance
Taus publicly states that tausdatabot fully respects robots.txt directives, including Disallow, Allow, and Crawl-Delay rules. In independent tests by webmasters (reported on various server logs), the bot was observed to obey disallowed paths immediately and to reduce request frequency when a Crawl-Delay was present. Taus’s own crawling policy page explicitly instructs site owners to use standard robots.txt entries to control access.
🔍 Detection Indicators
The primary User-Agent string is tausdatabot/1.0 (compatible; +https://taus.com/crawler). In addition, the bot includes the custom header X-Taus-Crawler: true and a From header containing the email address [email protected]. Some server logs also report a User-Agent variant of Mozilla/5.0 (compatible; TausDataBot/1.0) for compatibility with strict servers. The IP ranges are not publicly listed on a single page, but Taus recommends verifying via DNS PTR records pointing to *.taus.com.
📊 Data Usage
Data collected by tausdatabot is ingested into Taus’s data refinement engine, where it is cleaned, deduplicated, and annotated for use in AI training datasets. Taus’s product page (taus.com/products) confirms that curated data from this crawler is sold to enterprise clients for fine-tuning LLMs, building RAG systems, and training custom models. The bot does not store personal or copyrighted material beyond fair-use aggregation, and Taus offers an opt-out mechanism via their website.
⚙️ Rate Limiting Policy
Taus explicitly permits site owners to rate-limit tausdatabot without blocking it entirely, recommending thresholds of 100 requests per hour per IP before restricting further. This policy balances the bot’s legitimate need for data with the host’s resource protection, following industry best practices for cooperative crawling.
Similar Threats
Free Traffic Analysis
What's Actually Crawling Your Website?
Discover which unwanted bots are being blocked on your site, how often they hit, and where they come from — real data from your own traffic, not guesswork.
🔍 Scan My Site FreePowered by JA4 fingerprinting, honeypot traps & behavioral analysis
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.