Skip to main content

Boteraser | Website and Server Security Solutions

wume_crawler

Crawler User-Agent: wume-crawler

🤖 Overview

wume_crawler is a web crawler operated by Wume Technology, a Chinese AI company, first publicly documented in early 2023. Its purpose is to collect publicly accessible web content for training large language models and improving Wume's AI-driven search and recommendation products, including the Wume Assistant. The crawler is registered on common robot exclusion lists and is recognized as a legitimate bot.

🌐 Technical Behavior

The crawler utilizes distributed IP ranges from ASN 138950 (Wume Technology) and ASN 37963 (Huawei Cloud), primarily located in mainland China. According to official documentation (docs.wume.ai/crawler), it sends a maximum of 10 requests per second per domain with a mandatory crawl delay of at least 1 second. It supports HTTP/1.1 and HTTP/2, gzip compression, and respects ETag and If-Modified-Since headers for efficient recrawling. The bot uses persistent connections with Connection: keep-alive and requests gzip encoding via Accept-Encoding: gzip. It respects the Crawl-Delay directive in robots.txt and automatically reduces its rate upon receiving HTTP 429 responses. It does not execute JavaScript or render dynamic content, focusing on static HTML pages. It also adheres to the Accept-Language header, defaulting to English when unspecified. Wume publishes its crawler IP ranges at https://wume.ai/crawler-ips.

📋 robots.txt Compliance

Wume Technology's official robots.txt policy (https://wume.ai/robots-txt-policy) confirms that wume_crawler fully honors Disallow directives, including wildcard patterns and crawl-delay instructions. A 2024 update fixed a previous inconsistency with noindex meta tags on subpages when parent directories were allowed. The bot does not crawl pages explicitly blocked by robots.txt without permission.

🔍 Detection Indicators

The primary User-Agent string is "Mozilla/5.0 (compatible; wume_crawler/1.0; +https://wume.ai/crawler)". Alternate strings include "WumeBot" and "wume-crawler" for older versions. All requests include a custom HTTP header "X-Wume-Visit: true" and a cookie named "wume_session" with a 30-minute TTL. Reverse DNS lookups of requesting IPs resolve to *.wume.ai or *.wume-cloud.com domains. The

53% of Web Traffic Is Bots in 2026

— Imperva Bad Bot Report 2026

How much of your traffic is automated? Get your personal bot traffic report and see exactly what's hitting your server — completely free.

📊 Get My Bot Report

Sign up in seconds  ·  No card required

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.