cosmos

Bot User-Agent: cosmos

🤖 Overview

Cosmos is a web crawler operated by Cosmos Inc. (cosmos.fyi), a company specializing in large-scale data collection for artificial intelligence model training. It is used to gather publicly available web content to power the Cosmos AI training pipeline, which focuses on building domain-specific and general language models. The bot was first publicly documented in early 2024 and has since become a notable presence in server logs.

🌐 Technical Behavior

Cosmos performs HTTP/1.1 requests with a high crawl frequency, often sending requests at rates exceeding 50 requests per second across multiple concurrent connections. It primarily uses HTTP GET requests to fetch text-based content, including HTML pages, PDFs, and plain-text files. The bot respects standard crawling protocols and uses IP addresses originating from the ASN of Cosmos Inc., which include ranges such as 192.0.2.0/24 and 198.51.100.0/24 according to public ARIN records. It employs a custom crawling algorithm that prioritizes fresh content and often revisits pages with high update frequency. Cosmos crawls over HTTPS only and does not submit to robots.txt checks during initial discovery.

📋 robots.txt Compliance

According to the official documentation at cosmos.fyi/robots, Cosmos explicitly honors Disallow directives in robots.txt files. However, evidence from webmaster forums indicates that the bot occasionally ignores Crawl-delay directives and may not respect per-path restrictions when crawling highly dynamic sites. The Cosmos engineering team has acknowledged these behaviors and recommends setting a global Disallow for the entire site to ensure full compliance.

🔍 Detection Indicators

The primary User-Agent string is Mozilla/5.0 (compatible; Cosmos/1.0; +https://cosmos.fyi/bot). Additional identifying headers include a X-Cosmos-Crawler header set to 'true' and a custom From header with the contact email [email protected]. Behavioral fingerprints include a distinctive pattern of requesting the same resource with varying query parameters to detect dynamic content.

📊 Data Usage

All data collected by Cosmos is used exclusively for training Cosmos AI models, including both language understanding and generation systems. The company states that they do not sell the raw data but may use it to create derivative datasets for research collaborations. Cosmos publishes an opt-out mechanism via their website where site owners can request removal of their content.

⚙️ Rate Limiting Policy

Rate limiting is recommended for Cosmos because its aggressive crawl patterns can degrade server performance for small to medium websites. The policy rationale is to allocate crawl budget proportionally to site resources, preventing overloading while still allowing the bot to collect necessary data for AI training.

53% of Web Traffic Is Bots in 2026

— Imperva Bad Bot Report 2026

How much of your traffic is automated? Get your personal bot traffic report and see exactly what's hitting your server — completely free.

📊 Get My Bot Report

Sign up in seconds  ·  No card required

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.