Skip to main content

Boteraser | Website and Server Security Solutions

wanadoo

Bot User-Agent: wanadoo

🤖 Overview

Wanadoo is a web crawler operated by the former French telecommunications company Wanadoo (now part of Orange S.A.), originally used to index web pages for the Wanadoo search engine and directory service launched in the late 1990s. According to historical documentation archived on the Internet Archive and references in the robotstxt.org database, the Wanadoo crawler was part of the company’s initiative to build a proprietary search index for French-language content, complementing its ISP and web hosting services. The bot is now largely deprecated but may still be encountered on legacy systems or in residual configurations.

🌐 Technical Behavior

The Wanadoo crawler typically follows standard HTTP/1.1 protocols and uses a request frequency of approximately one request per second during a crawl session, based on archived logs. Its IP ranges historically belong to French address blocks assigned to Wanadoo’s backbone (e.g., 80.12.0.0/16 as per RIPE NCC records). The bot performs depth-first traversal of hyperlinks and respects standard crawl delays if specified in a site’s Crawl-Delay directive. It does not support modern protocols such as HTTP/2 or TLS 1.3, and its requests often lack a Referer header. The crawler’s behavior is non‑aggressive but can occasionally burst during initial indexing of large sites.

📋 robots.txt Compliance

Based on archived examples of robots.txt files from the early 2000s and the robots.txt specification maintained by Google, the Wanadoo crawler fully honors Disallow directives. It also respects Allow rules when present, as documented in the Web Robots Pages (robotstxt.org). No reports of deliberate violation of robots.txt exist in public security advisories or CVE entries.

🔍 Detection Indicators

The primary user‑agent string is Wanadoo (crawler) or simply Wanadoo, sometimes with a version suffix such as Wanadoo/1.0. Behavioral fingerprints include a consistent Accept header of text/html, application/xhtml+xml and the absence of a From or User-Agent override. Reverse DNS lookups on requesting IPs often resolve to subdomains like crawl.wanadoo.fr or proxy.wanadoo.fr.

📊 Data Usage

Collected web pages were used to build a searchable index for Wanadoo’s portal (later rebranded as Voila and eventually integrated into Orange’s services). The data was not employed for AI training or machine‑learning models, as the crawler predates modern LLM development. Usage was strictly for search indexing and directory listing of French websites.

⚙️ Rate Limiting Policy

Because the Wanadoo crawler does not implement dynamic adaptive rate control and may occasionally rescan entire site sections without caching, it is rate‑limited to prevent unnecessary load on servers. A threshold of 100 requests per 60 seconds from the same IP block is a common security best practice, ensuring resources are reserved for human users while still allowing legitimate indexing.

⚠️

Your Site May Be Hemorrhaging Revenue to Bots

Unwanted bots inflate your analytics, drain server resources, and slow down real users. Check if your site is affected — completely free.

Check My Site for Free

Free to start  ·  Cancel anytime

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.