webcopy
Bot User-Agent:webcopy
🤖 Overview
WebCopy is a legitimate desktop and command-line application developed by Cyotek Ltd., first released in 2012, designed for offline website mirroring and content archiving. It is operated by individual users, organizations, or researchers who need to preserve web pages for local browsing, compliance audits, or digital preservation. The primary product is the Cyotek WebCopy tool (available at https://www.cyotek.com/cyotek-webcopy), which crawls websites to copy static and dynamic assets (HTML, CSS, JavaScript, images, etc.) while maintaining relative link structures. Unlike large-scale search engine bots, WebCopy is a client-side agent that mimics a browser to fetch resources, and its usage is driven by manual triggers rather than continuous autonomous crawling.
🌐 Technical Behavior
WebCopy performs breadth-first or depth-first crawling based on user configuration, with a default concurrency level of 5 simultaneous connections (adjustable in advanced settings). It respects HTTP headers such as Last-Modified and ETag to avoid redundant downloads, and it intelligently throttles requests to prevent overloading servers — the default delay between requests is 250ms (configurable via the Request Delay option). The crawler does not emit a fixed IP range; instead, it uses the IP address of the machine running the application, which can be any residential or business IP. The tool supports HTTPS, HTTP, and FTP protocols, and can handle cookie-based sessions and form authentication when pre-configured. Importantly, WebCopy does not send a unique user‑agent string by default; it impersonates common browsers (e.g., Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0) to avoid blocking, though users can customize this via settings. The official documentation (https://www.cyotek.com/cyotek-webcopy/docs) explicitly warns operators to set appropriate delays and respect server load.
📋 robots.txt Compliance
WebCopy does not automatically parse or honor robots.txt directives by default, as noted in the Cyotek WebCopy FAQ (https://www.cyotek.com/cyotek-webcopy/faq). The developers state that the tool “does not read robots.txt because it is designed for offline browsing of public content under manual user control.” However, users can manually configure exclusion filters (e.g., “Block URLs matching */admin/*”) to simulate compliance. This means administrators cannot rely on robots.txt to block WebCopy — they must detect and rate-limit it via custom user‑agent strings or behavioral patterns.
🔍 Detection Indicators
WebCopy can be identified by its tendency to request a high volume of resources (CSS, JS, images) in rapid succession, often within seconds of each other, and by the absence of a consistent user‑agent string. When configured with the default settings, it sends a generic browser user‑agent, but advanced users may set a custom string like “WebCopy/1.0” (mentioned in community forums). Another fingerprint is the lack of Accept-Language and Referer headers in many requests, or the presence of unusual Connection: keep-alive patterns. Cyotek recommends that server operators monitor for repeated requests to the same domain with identical Accept headers and rapid timing.
📊 Data Usage
The collected data is used exclusively for local offline access — end users store the mirrored website on their own machines for archival, research, or compliance purposes. No data is transmitted to Cyotek or third parties. Because the tool is client‑side, any further use (e.g., republishing or training AI models) is determined solely by the operator and not by the software itself. In corporate environments, WebCopy is often deployed to preserve internal documentation or to create snapshots of regulatory‑required public content.
⚙️ Rate Limiting Policy
Although WebCopy is legitimate, its aggressive default concurrency (5 requests per second without built‑in throttling) can overwhelm smaller servers. Therefore, rate limiting is essential — thresholds such as 20 requests per 10 seconds per IP with a temporary 429 response are recommended. The policy rationale is to protect server resources from inadvertent denial‑of‑service caused by a single overzealous client while allowing legitimate human visitors and search engine bots unrestricted access.
🛡️
Stop Bots. Save Bandwidth. Protect Revenue.
Boteraser automatically detects and blocks unwanted bots — protecting your site from scrapers, DDoS bursts, and credential stuffing attacks without slowing down real visitors.
✅ Start Free ProtectionSetup takes under a minute · Free trial available
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.