clushbot
Bot User-Agent:clushbot
🤖 Overview
ClushBot is a web crawler operated by Clush Inc., a South Korean artificial intelligence company founded in 2020 and headquartered in Seoul. The bot began crawling in early 2022 and serves to index public web content for Clush's proprietary search engine and to collect training data for their large language models, including the ClushGPT series. The bot is integral to building their AI knowledge graph.
🌐 Technical Behavior
ClushBot performs both broad and targeted crawls using a multi-threaded engine that defaults to maintaining 4 simultaneous connections per domain. It respects the Crawl-Delay directive in robots.txt, with a default interval of 5 seconds between requests. The crawler originates from IP address ranges within South Korean netblocks (121.78.x.x and 14.63.x.x) and also utilizes cloud infrastructure from AWS (us-east-1 and ap-northeast-2) and Cloudflare (AS13335). ClushBot makes standard HTTP GET requests using HTTP/1.1, prefers HTTPS, and does not execute JavaScript or fetch embedded resources like images. It identifies itself via the User-Agent header and includes a From header with the email [email protected] for site owner inquiries.
📋 robots.txt Compliance
ClushBot fully complies with the Robots Exclusion Protocol. According to their official documentation at clush.com/robots, the bot checks for a robots.txt file before every crawl session and obeys both Disallow and Crawl-Delay directives. Independent site administrator reports confirm that ClushBot respects noindex meta tags and nofollow link attributes, with no evidence of bypassing these controls.
🔍 Detection Indicators
The primary User-Agent string is ClushBot/1.0 (compatible; ClushBot/1.0; +https://clush.com/bot). Additional variants include ClushBot/2.0 and ClushBot-Mobile/1.0 for mobile-specific crawls. Behavioral fingerprints include a consistent crawl rate of 10-15 requests per minute, a low concurrency of 2-4 simultaneous connections, and a custom request header X-Clush-Crawl: 1. The bot's reverse DNS often resolves to crawler.clush.com.
📊 Data Usage
Data collected by ClushBot serves two primary purposes: enhancing Clush's search index relevance and training their generative AI models. Raw HTML text is parsed, deduplicated, and stored in a vector database for retrieval-augmented generation (RAG). Structured data such as schema.org markup is extracted to improve entity recognition. Clush's privacy policy confirms that personal information is filtered out during processing and that collected data is not sold to third parties.
⚙️ Rate Limiting Policy
ClushBot is rate-limited because its aggressive, multi-threaded crawling can consume significant bandwidth on large sites with deep link structures. A threshold-based policy limiting requests to 20 per minute per IP is recommended to protect server resources while still allowing the bot to index content effectively.
Similar Threats
⚠️
Your Site May Be Hemorrhaging Revenue to Bots
Unwanted bots inflate your analytics, drain server resources, and slow down real users. Check if your site is affected — completely free.
Check My Site for FreeFree to start · Cancel anytime
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.