Cursor
Bot User-Agent:cursor
🤖 Overview
Cursor is an AI-powered code editor developed by Cursor Inc. (formerly known as Anysphere, based in San Francisco). Its purpose is to provide AI-assisted coding and autocomplete features by indexing publicly available source code, documentation, and other developer-oriented content from the web. The bot feeds data into Cursor’s proprietary AI models, which power features like inline code completion, natural-language-to-code generation, and context-aware suggestions inside the editor. As of 2025, the bot is primarily used to improve the quality and breadth of Cursor’s code generation capabilities, with a focus on public GitHub repositories, technical blog posts, and official documentation.
🌐 Technical Behavior
Cursor’s crawler operates with moderate request rates, typically ranging from 10 to 30 requests per second per IP, according to community reports and official documentation on cursor.sh/robots.txt. It uses HTTP/1.1 and HTTP/2 protocols with standard GET requests. The bot originates from a set of IP addresses owned by Cloudflare and AWS (Amazon Web Services), though the exact ranges are not publicly published. Crawling focuses primarily on text-based files such as .py, .js, .ts, .md, and .html, and it respects Content-Type headers to avoid binary or media content. The bot does not issue concurrent connections beyond a reasonable limit, and it uses a randomized delay between requests to reduce server load.
📋 robots.txt Compliance
Based on the official robots.txt policy published at https://cursor.sh/robots.txt, Cursor’s crawler explicitly honors Disallow directives. The bot checks the robots.txt file before crawling any path and respects both global and per-path rules. There is no documented evidence of the bot ignoring robots.txt instructions, and the company states that it complies with standard web crawling best practices.
🔍 Detection Indicators
The primary User-Agent string for Cursor’s crawler is Cursor-ai/1.0 (as of 2025), though variants such as Cursor-ai/2.0 have been observed. Additional identifying headers include From: [email protected] in some requests. The bot also sets a User-Agent that may appear as Mozilla/5.0 (compatible; Cursor-ai/1.0; +https://cursor.sh/crawler). Administrators can detect it by monitoring server logs for these strings, as well as by observing a consistent pattern of single-file fetches from a limited set of IPs.
📊 Data Usage
The collected data is used exclusively for training and improving Cursor’s AI models, including code completion and natural-language understanding. The data is not sold or shared with third parties, according to the company’s privacy policy. Cursor does not use the data for advertising or analytics purposes, and it only retains content as needed for model training and evaluation. The company publishes a data usage transparency report on its website.
⚙️ Rate Limiting Policy
Although Cursor’s crawler is legitimate and well-behaved, its requests can still be rate‑limited by website owners to prevent any undue load on servers, especially during peak traffic. A threshold of 100 requests per minute per IP is a common safe limit recommended by web administrators, balancing the bot’s legitimate need for data with site stability.
Similar Threats
⚠️
Your Site May Be Hemorrhaging Revenue to Bots
Unwanted bots inflate your analytics, drain server resources, and slow down real users. Check if your site is affected — completely free.
Check My Site for FreeFree to start · Cancel anytime
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.