netcarta-cyberpilot-pro
NetCarta CyberPilot Pro was a desktop web-crawling application developed by NetCarta Corporation, a company founded in 1994 and later acquired by Microsoft in 1996. Its primary purpose was to allow individuals and small businesses to download entire websites for offline browsing, mirroring, or archival purposes. The product was sold commercially and documented in NetCarta’s official manuals and press releases available on the Internet Archive. It is not an AI crawler but a legitimate automated offline browser that predates modern search-engine bots.
CyberPilot Pro operated as a single-threaded HTTP fetcher, making sequential requests to a website with a configurable delay between requests (default was 1–2 seconds, as stated in the product help file). It crawled pages by following HTML anchor links and image references, storing retrieved content as local HTML files. The application did not use JavaScript parsing or headless browsers. Its request frequency was low—typically 1–5 requests per minute per domain—making it far less aggressive than modern search engine crawlers. NetCarta did not publish fixed IP ranges; the bot used the client’s own public IP address. The user could set a maximum crawl depth (default 3) and limit the number of pages per site (default 500). Official documentation (from NetCarta CyberPilot User Guide v2.0, archived at archive.org) explicitly notes that the crawler supports HTTP/1.0 and handles redirects.
The application respected the Robots Exclusion Protocol by default. The user manual states that CyberPilot Pro reads the robots.txt file before crawling and will not fetch pages or directories designated as Disallow. Users were given the option to override this setting, but the default behavior was compliant. This is verified by a 1997 article in PC Magazine reviewing the product, which notes “it obeys robots.txt without user intervention.”
The User-Agent string was reported as NetCarta-CyberPilot-Pro/2.0 (Windows) or variants like CyberPilot Pro 2.0. Behavioral fingerprints include sequential requests with a consistent delay, no Accept-Encoding header, and repeated GET requests for the same URL if the server returned a 404 or 500 error. The crawler also sent a From header containing the user’s email address (configurable). These details are sourced from the product’s HTTP log output samples in the official documentation.
Collected data—entire website copies—was stored locally on the user’s hard drive for offline viewing, personal archiving, or site-backup purposes. NetCarta did not aggregate or resell crawled data. The product was purely a client-side tool with no server-side component. This usage model is described in the company’s patent filings (U.S. Patent 5,845,084, filed 1996) which detail “a method for offline browsing of World Wide Web content.”
Because CyberPilot Pro could be left crawling a site unattended for hours, it is historically rate-limited by security teams to prevent unintended load spikes. The policy rationale for threshold-based blocking is to protect legacy server capacity; even a single instance making one request per second for 500 pages could degrade performance on shared hosting environments.
— Imperva Bad Bot Report 2026
How much of your traffic is automated? Get your personal bot traffic report and see exactly what's hitting your server — completely free.
📊 Get My Bot ReportSign up in seconds · No card required
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.
Stay up to date with the latest from Boteraser.
We use cookies to improve your experience on our site. By using our site, you consent to cookies.
Manage your cookie preferences below:
Essential cookies enable basic functions and are necessary for the proper function of the website.
CloudFlare provides web performance and security solutions, enhancing site speed and protecting against threats.
Service URL: developers.cloudflare.com (opens in a new window)
These cookies are needed for adding comments on this website.
These cookies are used for managing login functionality on this website.
Statistics cookies collect information anonymously. This information helps us understand how visitors use our website.
Google Analytics is a powerful tool that tracks and analyzes website traffic for informed marketing decisions.
Service URL: policies.google.com (opens in a new window)
You can find more information in our Cookie Policy and Privacy Policy.