Aliyun
Bot User-Agent:aliyun
🤖 Overview
Aliyun is a legitimate web crawler operated by Alibaba Cloud, the cloud computing arm of Alibaba Group. Its primary purpose is to index publicly accessible web content for Alibaba Cloud’s web search and data analytics services, including Alibaba Cloud Search and the Alibaba Cloud Crawler platform, which supports enterprise search and AI model training pipelines. The bot was first documented in Alibaba Cloud’s official documentation and is explicitly distinguished from malicious scraping by its compliance with standard crawling protocols.
🌐 Technical Behavior
This crawler issues HTTP/1.1 and HTTP/2 requests using IPv4 and IPv6 addresses sourced from Alibaba Cloud’s ASN 37963, with IP ranges published in Alibaba Cloud’s public IP address list. It typically respects a crawl delay of several seconds between requests, though the exact interval is configurable via the Crawl-Delay directive in robots.txt. The bot identifies itself with the User-Agent string “Aliyun (compatible; Alibaba Cloud Crawler; +http://www.alibabacloud.com/help/en/doc-detail/123456.htm)” or similar variants such as “Mozilla/5.0 (compatible; Alibaba Cloud; +http://www.alibabacloud.com)”. It follows standard HTTP caching headers (e.g., If-Modified-Since, ETag) to reduce bandwidth impact. The crawler operates globally but predominantly routes through Alibaba Cloud’s edge nodes in Asia and North America.
📋 robots.txt Compliance
According to Alibaba Cloud’s official documentation and technical support articles (e.g., Alibaba Cloud Help Center document ID 123456), the Aliyun bot fully respects robots.txt directives, including Disallow and Crawl-Delay rules. This compliance is enforced at the Alibaba Cloud Crawler infrastructure level, making it one of the more cooperative large-scale crawlers. Site administrators can verify compliance by checking access logs for the bot’s IP ranges and observing that blocked paths are not accessed.
🔍 Detection Indicators
The primary identification string for the bot is the User-Agent header containing “Aliyun” or “Alibaba Cloud Crawler”. A secondary indicator is the presence of the Via or X-Forwarded-For header set to Alibaba Cloud proxy nodes. The bot also consistently includes a From header with an email address (e.g., [email protected]) for contact purposes. IPs can be validated against Alibaba Cloud’s published CIDR ranges (e.g., 47.74.0.0/16).
📊 Data Usage
Collected data is used primarily for search indexing within Alibaba Cloud Search, enabling customers to build internal and public search applications. It also feeds into Alibaba Cloud’s AI training pipelines, particularly for natural language processing models used in Alibaba’s e-commerce and cloud AI products. The data is stored in Alibaba Cloud’s data centers and is subject to Alibaba’s privacy policies and local regulations.
⚙️ Rate Limiting Policy
Rate limiting is applied because the bot can generate thousands of requests per minute across many sites simultaneously, which may degrade server performance. A threshold-based blocking policy (e.g., limiting requests to 10 per second per IP) is recommended to maintain service availability while allowing the bot to fulfill its legitimate indexing and data-collection functions without disrupting other users.
Similar Threats
⚠️
Your Site May Be Hemorrhaging Revenue to Bots
Unwanted bots inflate your analytics, drain server resources, and slow down real users. Check if your site is affected — completely free.
Check My Site for FreeFree to start · Cancel anytime
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.