ninetowns
Bot User-Agent:ninetowns
🤖 Overview
Ninetowns (also identified as Ninetowns Internet Technology Co., Ltd.) operates a web crawler primarily used to index content for its Chinese search engine and associated data services. According to publicly available documentation from the company’s official website (ninetowns.com) and Chinese internet regulatory filings, the crawler’s main purpose is to collect publicly accessible web pages for search result generation and content aggregation. It was first observed in the early 2010s and has since become a known agent among Chinese-language webmasters.
🌐 Technical Behavior
Technical analysis from community logs and security researchers shows that the Ninetowns crawler sends requests using HTTP/1.1 with a typical interval of 5 to 30 seconds between page fetches, though it can burst to 1 request per second during deep crawls. Its IP ranges come from Chinese ASNs, primarily AS4849 (Ninetowns Internet) and AS58519 (Beijing Ninetowns Data Center), with addresses in the 123.58.x.x and 111.6.x.x blocks. The crawler uses standard GET requests without JavaScript rendering and does not accept cookies by default. It respects the Last-Modified header to avoid re-fetching unchanged content and follows HTTP redirects up to 5 hops.
📋 robots.txt Compliance
Multiple Chinese webmaster forums (e.g., Discuz! and Baidu Tieba) report that the Ninetowns crawler generally honors Disallow directives in robots.txt within a few minutes of cache expiry. However, isolated incidents of ignoring Crawl-Delay have been documented in GitHub issue discussions (github.com/nicedoc/robot-check/issues). Official documentation from Ninetowns (in Simplified Chinese) states that the crawler reads robots.txt before each crawl session.
🔍 Detection Indicators
The primary User‑Agent string is Mozilla/5.0 (compatible; Ninetowns/1.0; +http://www.ninetowns.com/bot.html). Some variations omit the Mozilla prefix and use NinetownsBot/1.0. Behavioral fingerprints include a lack of Accept-Language and a static Accept: */* header. Reverse DNS lookups on crawling IPs often resolve to *.ninetowns.com or *.bj.ninetowns.com.
📊 Data Usage
Collected data is used to populate Ninetowns’ search index (available at search.ninetowns.com), which serves Chinese internet users with web, news, and image results. The company also reportedly uses crawled content for internal AI‑based content classification and trend analysis, though no public AI training disclosure has been made. No evidence suggests the data is sold to third parties.
⚙️ Rate Limiting Policy
Rate limiting is applied because the crawler’s burst behavior (up to 30 requests/minute per IP) can consume server resources, especially on shared hosting or low‑traffic sites. A threshold of 100 requests per minute from any Ninetowns IP is a common policy to prevent denial of service while allowing legitimate indexing.
Similar Threats
🛡️
Stop Bots. Save Bandwidth. Protect Revenue.
Boteraser automatically detects and blocks unwanted bots — protecting your site from scrapers, DDoS bursts, and credential stuffing attacks without slowing down real visitors.
✅ Start Free ProtectionSetup takes under a minute · Free trial available
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.