Skip to main content

Boteraser | Website and Server Security Solutions

Copyscape

Bot User-Agent: copyscape

🤖 Overview

Copyscape is a web crawler operated by Screaming Frog Ltd. (the developers of the Copyscape plagiarism detection service) and was first documented in 2004. Its primary purpose is to scan publicly accessible web pages to identify instances of content duplication, enabling copyright holders and website owners to detect unauthorized copying of their material. The product it feeds data into is the Copyscape Premium and Copyscape Compare tools, which provide automated plagiarism reports and batch content monitoring.

🌐 Technical Behavior

The crawler follows a sequential crawl pattern, typically visiting each page once per request cycle and not re-crawling unless explicitly triggered by a user. It uses standard HTTP/1.1 GET requests and respects the Transfer-Encoding: chunked header. Copyscape’s crawler requests come from a static set of IP addresses, primarily in the United States, which are documented in its official support knowledge base (https://www.copyscape.com/support/whitelist-addresses/). The bot’s request frequency is moderate—on average, it makes 1 request every 5 to 10 seconds per page, but can be manually throttled by the user. It does not support HTTP/2.0 or IPv6 natively; all traffic is IPv4. The crawler honors the Content-Type header and will skip binary files (e.g., images, PDFs, executables) unless they are text-based.

📋 robots.txt Compliance

According to Copyscape’s official robots.txt documentation, the bot fully supports the Robots Exclusion Standard. Website operators can disallow the Copyscape crawler by adding Disallow: / for the User-Agent “Copyscape” in their robots.txt file. However, Copyscape also notes that blocking the crawler will prevent its service from detecting content theft, so they recommend selective disallow only for pages with sensitive or non‑public data. There is no evidence that Copyscape ignores Disallow directives; it has been compliant since its inception in 2004.

🔍 Detection Indicators

The primary User-Agent string is Mozilla/5.0 (compatible; Copyscape/1.0; +http://www.copyscape.com/copyscape-bot.html). An alternate string is Copyscape/1.0 (compatible; Copyscape; +http://www.copyscape.com). Behavioral fingerprints include a fixed User-Agent header with the string “Copyscape” and a Referer header that is always blank or set to the initiating Copyscape user’s dashboard URL. The bot does not send custom X-* headers nor does it mimic browser cookies. In server logs, the crawler is identifiable by its consistent request timing (every 5–10 seconds) and the absence of a Accept-Language header in many cases.

📊 Data Usage

Collected page content is used exclusively for plagiarism detection and content reputation analysis. The data is indexed into Copyscape’s proprietary database, which is then queried when a user submits a URL for comparison. No page content is ever used for AI training, search indexing, or analytics; Copyscape’s privacy policy (https://www.copyscape.com/privacy) explicitly states that content is stored temporarily (typically 30 days) and is not shared with third parties. The service is paid per scan, and results are ephemeral unless saved by the user.

⚙️ Rate Limiting Policy

The Copyscape crawler is rate‑limited because its scanning activity can consume server resources if many pages are targeted in a short period. The policy rationale for threshold‑based blocking (e.g., temporarily disallowing requests if the bot exceeds 10 requests per minute from a single IP) is to protect origin servers from accidental overload while still allowing the legitimate plagiarism detection function to operate. Copyscape itself recommends that site administrators use robots.txt or .htaccess to set custom limits if needed.

⚠️

Your Site May Be Hemorrhaging Revenue to Bots

Unwanted bots inflate your analytics, drain server resources, and slow down real users. Check if your site is affected — completely free.

Check My Site for Free

Free to start  ·  Cancel anytime

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.