webdatacentrebot
Bot User-Agent:webdatacentrebot
🤖 Overview
webdatacentrebot is a web crawler operated by Web Data Centre Ltd, a UK-based company specializing in web scraping and data aggregation services for e-commerce, price monitoring, and market research. First publicly documented around 2016, the bot systematically indexes publicly accessible web pages to populate the company’s proprietary data products, including competitive pricing feeds and product catalog databases. Unlike general-purpose search engine bots, webdatacentrebot is explicitly designed for commercial data extraction rather than search index building.
🌐 Technical Behavior
The crawler follows a breadth-first traversal pattern, issuing requests at an average rate of one to five requests per second per source IP address, according to operational logs shared on community forums and the official webdatacentre.com documentation. It supports both HTTP/1.1 and HTTP/2 protocols and honours the If-Modified-Since header to reduce redundant transfers. IP ranges used by webdatacentrebot are drawn from the 78.129.200.0/21 block (AS12488, Krystal Hosting) and 185.151.24.0/22 (AS20860, Iomart Hosting), as verified through reverse DNS lookups and WHOIS records published by RIPE NCC. The crawler does not execute JavaScript or render pages; it processes only static HTML content and linked CSS/JavaScript resources necessary for layout parsing, as stated in the company’s technical FAQ. Concurrent connections are limited to two per host by default, though this can increase during scheduled deep crawls.
📋 robots.txt Compliance
Web Data Centre explicitly states in its Crawler Policy page (webdatacentre.com/crawler-policy) that webdatacentrebot fully respects Disallow directives in robots.txt files, and advises site owners to use the standard User-agent: webdatacentrebot syntax. Archived correspondence from the company’s support team on WebmasterWorld confirms that the bot checks robots.txt at the beginning of each crawl session and caches the parsed rules for the session’s duration.
🔍 Detection Indicators
The canonical User-Agent string is Mozilla/5.0 (compatible; webdatacentrebot/2.0; +https://webdatacentre.com/bot.html), though older versions may appear as webdatacentrebot/1.0. A secondary string, WebDataCentreBot/2.0, has been observed in server logs from 2022 onward. Behavioral fingerprints include a request interval of 0.8–1.2 seconds and a strict preference for port 80/443 with no cookie persistence. Requests lack typical browser headers such as Accept-Language or Upgrade-Insecure-Requests.
📊 Data Usage
Collected data is fed into Web Data Centre’s commercial product suite, including real-time price comparison APIs, historical pricing dashboards, and product attribute enrichment services. The company’s privacy policy (webdatacentre.com/privacy) clarifies that no personal identifiable information is intentionally harvested; only publicly available content related to products, services, and public contact details is retained. Aggregated datasets are sold to e-commerce platforms and analytics firms for market trend analysis.
⚙️ Rate Limiting Policy
Webdatacentrebot is rate-limited on many production sites because its sustained crawl volume can degrade server performance during peak hours, especially on shared hosting environments. A threshold-based block is justified by the bot’s commercial, non-essential nature—unlike search engine crawlers that provide organic traffic, webdatacentrebot extracts value without driving visitors, making reasonable rate limiting a standard practice to protect site resources.
Similar Threats
🛡️
Stop Bots. Save Bandwidth. Protect Revenue.
Boteraser automatically detects and blocks unwanted bots — protecting your site from scrapers, DDoS bursts, and credential stuffing attacks without slowing down real visitors.
✅ Start Free ProtectionSetup takes under a minute · Free trial available
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.