cd-preload

Bot User-Agent: cd-preload

🤖 Overview

The cd-preload bot is operated by CDNetworks, a global content delivery network (CDN) provider headquartered in South Korea. Its primary purpose is to preload or “warm” website caches by crawling URLs that are expected to receive high traffic, ensuring that cached copies are ready in CDNetworks edge servers before users request them. This proactive caching improves page load times and reduces origin server load, particularly during traffic spikes or scheduled content updates. The bot feeds data directly into CDNetworks’ caching infrastructure, not into external search engines or AI models.

🌐 Technical Behavior

The cd-preload crawler issues HTTP GET requests at variable rates depending on the configuration set by the CDNetworks customer. Typically, it sends one request every few seconds for each preloaded URL, but the frequency can be adjusted via the CDNetworks control panel. It uses standard HTTP/1.1 and HTTP/2 protocols and supports both IPv4 and IPv6. The IP addresses used by cd-preload belong to CDNetworks’ own ASN (AS38197) and are publicly listed in their official IP ranges documentation at https://www.cdnetworks.com/support/ip-ranges. The bot does not execute JavaScript, parse CSS, or follow redirects beyond two hops; it only fetches the exact URLs specified in the preload configuration. It sends a User-Agent header with the string “cd-preload” and may include a Via header referencing CDNetworks edge nodes. The default crawl pattern mimics a desktop browser, though it does not download images or other subresources unless explicitly listed.

📋 robots.txt Compliance

According to CDNetworks’ official documentation at https://www.cdnetworks.com/support/crawler-policies, the cd-preload bot fully respects robots.txt directives. It will not crawl any URL path disallowed by the site’s robots.txt file, and it checks the file before each crawl session. CDNetworks recommends that webmasters who wish to block cd-preload add Disallow: / for User-agent: cd-preload in their robots.txt, and the bot will honor that rule without exception.

🔍 Detection Indicators

The primary detection string is the User-Agent: Mozilla/5.0 (compatible; cd-preload; +http://www.cdnetworks.com/cd-preload.html). Additionally, the bot’s requests originate from IP addresses in CDNetworks’ owned ranges, which can be verified via WHOIS lookups or the published IP list. The bot does not set a custom X-Forwarded-For header by default, but may include a Via header indicating a CDNetworks proxy. Behavioral fingerprints include a fixed Accept header (*/*), no Accept-Language, and a consistent crawl interval configured by the customer.

📊 Data Usage

All data collected by cd-preload is used exclusively for cache warming within CDNetworks’ CDN platform. The bot fetches full HTML (and optionally CSS/JavaScript if specified) to store as complete cached objects on edge servers. No user data is collected, and the content is not shared with third parties or used for analytics, AI training, or search indexing. The cached files are retained per CDNetworks’ standard TTL settings and purged when the customer updates the preload configuration.

⚙️ Rate Limiting Policy

Rate limiting of cd-preload is recommended because its deliberate, persistent crawling can consume server resources, especially if many URLs are preloaded concurrently. Most web application firewalls (WAFs) set a threshold of 5–10 requests per second from a single IP to prevent unintended load, while still allowing the bot to complete its warmup task within a reasonable timeframe. This balance protects the origin server without blocking a legitimate preloading service.

🛡️

Stop Bots. Save Bandwidth. Protect Revenue.

Boteraser automatically detects and blocks unwanted bots — protecting your site from scrapers, DDoS bursts, and credential stuffing attacks without slowing down real visitors.

✅ Start Free Protection

Setup takes under a minute  ·  Free trial available

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.