crescent

Bot User-Agent: crescent

🤖 Overview

Crescent is a web crawler operated by Crescent AI Inc., a privately held company based in San Francisco, California, focused on building large-scale language models for enterprise AI applications. According to official documentation available at https://crescent.ai, the bot was first publicly identified in early 2024 and is designed to collect publicly accessible web content for training Crescent’s proprietary LLM, known as Crescent-7B and Crescent-70B. The product is a general-purpose natural language understanding platform used in customer support automation and content generation.

🌐 Technical Behavior

Crescent uses HTTP/1.1 with persistent connections (Keep-Alive) and respects standard ETag and If-Modified-Since headers to minimize redundant downloads. The crawler issues requests from a distributed pool of IP addresses primarily in the 104.16.0.0/12 and 172.64.0.0/13 ranges, which are owned by Cloudflare and used as a proxy layer for rate management, as detailed in Crescent’s IP list at https://crescent.ai/ips.txt. Crawl frequency is set to a default delay of 10 seconds between requests to the same host, but the bot can burst up to 5 concurrent requests per domain when authorized. It strictly follows the robots.txt directives, including Crawl-Delay, and supports the X-Robots-Tag HTTP header for page-level instructions. Crescent also uses Accept-Encoding: gzip and Accept-Language: en-US,en;q=0.9 to mimic modern browsers.

📋 robots.txt Compliance

According to Crescent’s publicly posted crawler policy on https://crescent.ai/robots-compliance, the bot fully honors Disallow directives in robots.txt and also supports the Allow directive for granular control. The company states that they parse the file at each visit and cache it for up to 24 hours, and they automatically respect Crawl-Delay values. Any violations of robots.txt are logged and reviewed; repeated non-compliance results in the IP being temporarily suspended from crawling that domain.

🔍 Detection Indicators

The primary User-Agent strings are "Mozilla/5.0 (compatible; Crescent/1.0; +https://crescent.ai)" and "CrescentBot/1.0". Behavioral fingerprints include a request rate no higher than 0.1 requests per second on average over 10-minute windows, and the use of a custom X-Crescent-Crawl-ID header containing a UUID. Reverse DNS lookups on crawling IPs resolve to crawler.crescent.ai or similar subdomains. The bot also sets a From header with the contact email [email protected] for compliance inquiries.

📊 Data Usage

Collected data is used exclusively for training and fine-tuning Crescent’s large language models, as described in their privacy policy at https://crescent.ai/privacy. The crawled content is preprocessed to remove personal identifiable information (PII) and is stored in an encrypted data lake. Crescent also uses the data to improve model safety, reduce bias, and power their enterprise AI search product, which summarizes web content for business intelligence.

⚙️ Rate Limiting Policy

Rate limiting is applied because Crescent, while legitimate, can generate significant traffic when crawling large sites, potentially degrading performance if left unchecked. The recommended policy, per Crescent’s own best practices guide, is to allow up to 10 requests per minute per IP and return HTTP 429 Too Many Requests with a Retry-After header of 60 seconds when exceeded, ensuring fair resource sharing across all crawlers.

Free Bot Analysis

Is Your Site Under Bot Attack Right Now?

Find out exactly how much of your traffic is automated — and which bots are draining your bandwidth and skewing your analytics.

Run Free Bot Scan →

No credit card required  ·  Results in minutes

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.