aidu

Bot User-Agent: aidu

🤖 Overview

Aidu (also referred to as AiduBot) is a legitimate web crawler operated by Beijing Aidu Technology Co., Ltd., a Chinese artificial intelligence company. According to the official documentation published at https://aidu.com/crawler, the bot’s purpose is to systematically collect publicly accessible web content for training and improving Aidu’s proprietary large language models (LLMs). The crawled data feeds directly into Aidu’s AI training pipeline, supporting natural language understanding and generation tasks. Unlike many search-engine crawlers, Aidu is exclusively focused on acquiring high-quality text and structured data for model refinement, not for indexing a public search engine.

🌐 Technical Behavior

AiduBot performs HTTP/1.1 and HTTP/2 GET requests with a default crawl frequency that is documented as “moderate” — typically between one request every 2–5 seconds per domain, though bursts may occur during initial site scans. The crawler respects robots.txt directives before every request and supports conditional GET using ETag and Last-Modified headers to reduce server load. IP ranges are not publicly listed but have been observed to originate from China-based ASNs such as AS4837 (China Unicom) and AS9808 (China Mobile). Aidu does not use any known proxy or anonymization services; all requests come from owned or leased datacenter IPs. The crawler can handle gzip and deflate compression, and its HTTP User-Agent header includes a version number (e.g., AiduBot/1.0).

📋 robots.txt Compliance

Based on the official robots.txt policy published at https://aidu.com/robots.txt, AiduBot explicitly honors the Disallow and Allow directives found in any website’s robots.txt file. The documentation states that the crawler reads robots.txt before every new crawl session and caches the file for up to 24 hours. Failure to fetch robots.txt does not cause the bot to proceed; instead, it stops crawling until a successful retrieval. Independent audits by webmasters confirm that AiduBot does not violate disallowed paths, aligning with its public commitment to ethical scraping.

🔍 Detection Indicators

The primary User-Agent string observed in server logs is AiduBot/1.0 (or Aidu/1.0 on older versions). Secondary identifiers include the From header containing [email protected] and a User-Agent suffix of +(https://aidu.com/bot) for verification. The bot also sends a specific HTTP request header Aidu-Version set to the crawler software revision. Reverse DNS lookups on connecting IPs often resolve to hostnames ending in .aidu.com, though this is not guaranteed.

📊 Data Usage

All text and metadata collected by AiduBot are used exclusively for training and evaluating Aidu’s large language models. As stated in the company’s data policy (https://aidu.com/data-policy), crawled pages are processed into tokenized datasets that are stripped of personally identifiable information and copyrighted content upon request. The data is not shared with third parties, resold, or used for advertising targeting. Aidu also maintains a opt-out form for webmasters who wish to exclude their content even if robots.txt is not set.

⚙️ Rate Limiting Policy

While AiduBot is a legitimate agent, it is rate-limited because its burst behavior can still overwhelm poorly configured servers. The recommended threshold-based blocking (e.g., 10 requests per second per IP) is a sensible policy to prevent unintended resource exhaustion without permanently denying access to a compliant crawler. Aidu’s documentation acknowledges that site operators may apply their own rate limits as long as they do not intentionally block the bot entirely.

53% of Web Traffic Is Bots in 2026

— Imperva Bad Bot Report 2026

How much of your traffic is automated? Get your personal bot traffic report and see exactly what's hitting your server — completely free.

📊 Get My Bot Report

Sign up in seconds  ·  No card required

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.