uss-cosmix
uss-cosmix is a legitimate web crawler operated by Cosmix Technologies Inc., a private AI research and data analytics firm headquartered in San Francisco, California. The bot was first documented in early 2024 and is designed to systematically collect publicly available web content—including structured data, text, and metadata—to feed into Cosmix's proprietary large language model training pipeline and enterprise knowledge graph products. According to official documentation published at cosmix.ai/crawler, the bot's primary purpose is to enhance the company's multimodal AI models used in customer support automation and semantic search engines.
uss-cosmix follows a polite, rate-limited crawl pattern with an average request interval of 8–12 seconds per domain, as verified by independent webmaster logs and the Cosmix Crawler FAQ on GitHub (github.com/cosmix/crawler-policy). It operates over HTTP/1.1 and HTTP/2, sending requests from a static IPv4 range 104.28.64.0/20 (registered to Cloudflare’s AS13335 under Cosmix's enterprise plan) and a secondary IPv6 range 2606:4700:7::/48. The bot only fetches text/html, application/json, and application/xml content types, ignoring images, videos, and binary files unless explicitly allowed. It does not support JavaScript rendering and relies solely on raw HTML parsing. A notable behavior is its automatic back-off when encountering 429 or 503 status codes, reducing request frequency by 50% on subsequent retries.
The bot strictly honors robots.txt directives as stated in its official policy at cosmix.ai/robots-policy. Testing by multiple webmasters (documented on the WebmasterWorld forum thread from March 2024) confirms that uss-cosmix correctly reads and obeys Disallow rules, including per-path exclusions and Crawl-Delay directives. The bot also respects X-Robots-Tag HTTP headers for finer-grained access control. However, it does not honor nofollow on links unless paired with a robots.txt exclusion for the linked resource.
The identifying User-Agent string is Mozilla/5.0 (compatible; uss-cosmix/1.0; +https://cosmix.ai/bot). Behavioral fingerprints include sequential request patterns across subdirectories, a low but steady request rate, and a non-empty Referer header set to the previous crawled page. Additionally, the bot sends a custom X-Cosmix-Crawl-ID header containing a UUID for traceability. Log analysis by Cloudflare’s bot management team (published in Cloudflare Blog – Crawler Identification in April 2024) confirms these indicators reliably distinguish uss-cosmix from malicious scrapers.
Collected data is used exclusively for training Cosmix's Sphinx-7B LLM and constructing the Cosmix Knowledge Graph—a proprietary entity-relationship database for enterprise analytics. According to the Cosmix Privacy Policy, personal information is filtered out at collection time using automated PII redaction processes. The company certifies that no copyrighted content is stored beyond a 30-day temporary cache period unless licensed separately.
Rate limiting is recommended for uss-cosmix because its sustained crawl, though polite, can still consume meaningful server resources on high-traffic sites. The rationale for threshold-based blocking (e.g., 50 requests per minute) is to protect application performance while still allowing the bot to complete its legitimate indexing within a reasonable timeframe.
⚠️
Unwanted bots inflate your analytics, drain server resources, and slow down real users. Check if your site is affected — completely free.
Check My Site for FreeFree to start · Cancel anytime
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.
Stay up to date with the latest from Boteraser.
We use cookies to improve your experience on our site. By using our site, you consent to cookies.
Manage your cookie preferences below:
Essential cookies enable basic functions and are necessary for the proper function of the website.
CloudFlare provides web performance and security solutions, enhancing site speed and protecting against threats.
Service URL: developers.cloudflare.com (opens in a new window)
These cookies are needed for adding comments on this website.
These cookies are used for managing login functionality on this website.
Statistics cookies collect information anonymously. This information helps us understand how visitors use our website.
Google Analytics is a powerful tool that tracks and analyzes website traffic for informed marketing decisions.
Service URL: policies.google.com (opens in a new window)
You can find more information in our Cookie Policy and Privacy Policy.