Claude-Web
Bot User-Agent:claude-web
🤖 Overview
Claude-Web is an automated web crawler operated by Anthropic, the AI company behind the Claude large language model family. Announced in a public support article published in August 2023, its primary purpose is to collect publicly accessible text content for training and improving Anthropic’s foundation models, including Claude 2 and subsequent versions. The crawler is designed to operate transparently and follows industry-standard protocols for legitimate data collection.
🌐 Technical Behavior
The Claude-Web crawler issues HTTP GET requests over protocols HTTP/1.1 and HTTP/2, targeting static content such as HTML documents and PDF files. It does not execute JavaScript, submit forms, or interact with dynamic elements. According to Anthropic’s official documentation, the crawler originates from IP ranges within ASN 200244, specifically subnets 20.96.0.0/16, 20.83.0.0/16, and 20.40.0.0/16. Each request includes a standard Accept header preferring text/html, and the crawler respects the crawler-level rate limit of approximately one request per second by default. It identifies itself via the User-Agent header Claude-Web/1.0 and includes a From header pointing to a contact email address for webmasters. Anthropic publishes its full IP range list and User-Agent string on their official support portal (support.anthropic.com) for easy access.
📋 robots.txt Compliance
Anthropic explicitly states that Claude-Web honors the robots.txt standard. Webmasters can block the crawler by adding a User-agent: Claude-Web directive to their robots.txt file, and the crawler will not access any disallowed paths. This compliance has been verified through multiple independent reports and is consistent with Anthropic’s responsible AI development policies.
🔍 Detection Indicators
The primary fingerprint is the User-Agent string Claude-Web/1.0. Additional indicators include the originating IP ranges (20.96.0.0/16, 20.83.0.0/16, 20.40.0.0/16) and the presence of a From header with the email address [email protected]. Reverse DNS records for crawling IPs resolve to hostnames under the anthropic.com domain, further verifying identity.
📊 Data Usage
All content collected by Claude-Web is used exclusively for training and improving Anthropic’s AI models, including the Claude series. The company states that personal identifiable information is stripped during processing, and the data is not used for advertising, resold, or shared with third parties. This aligns with Anthropic’s published commitment to ethical AI development and transparency.
⚙️ Rate Limiting Policy
Although Claude-Web is a legitimate, rate‑limited crawler, its sustained crawl campaigns can still consume significant server resources. Rate limiting is recommended to protect web application stability—typically thresholds of 10 requests per second per IP are effective—ensuring fair access for all users while still allowing Anthropic to collect necessary training data.
Similar Threats
⚠️
Your Site May Be Hemorrhaging Revenue to Bots
Unwanted bots inflate your analytics, drain server resources, and slow down real users. Check if your site is affected — completely free.
Check My Site for FreeFree to start · Cancel anytime
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.