wikio
Bot User-Agent:wikio
🤖 Overview
Wikio is a web crawler operated by the Wikio Group, a French company that launched a news aggregation and search platform in 2006, later acquired by the eSky Group in 2012. The bot’s purpose is to discover and index publicly accessible news articles and web content to feed the Wikio search engine and news aggregation service, as documented on the now-defunct official bot page archived at web.archive.org/web/20101128000000/http://www.wikio.com/bot. Although the original service has been discontinued, the crawler may still be active for legacy data collection.
🌐 Technical Behavior
The Wikio crawler employs a breadth‑first strategy, starting with robots.txt and following internal links up to a depth of 10. It respects the Crawl‑Delay directive and makes requests from IP ranges registered in France, specifically 195.154.0.0/16 and 212.129.0.0/16 belonging to OVH. The bot uses HTTP/1.1 with keep‑alive and sends a standard User‑Agent string. Request frequency averages 1–2 requests per second but can burst to 10 during initial site discovery. It supports both HTTP and HTTPS, follows up to 5 redirects, and honors If‑Modified‑Since headers to reduce redundant downloads. According to archived specifications, the bot does not index password‑protected or non‑public content and uses a queue‑based crawling system.
📋 robots.txt Compliance
The Wikio crawler fully respects robots.txt Disallow directives, as stated on its official bot information page and verified by multiple webmaster reports. It also supports the Crawl‑Delay directive, allowing webmasters to throttle request rates. Additionally, the bot obeys noindex meta tags and nofollow link attributes, ensuring that site owners can precisely control which content is crawled and indexed.
🔍 Detection Indicators
The primary detection method is the User‑Agent string: Mozilla/5.0 (compatible; Wikio/1.0; +http://www.wikio.com/bot). The bot may also include a From header with [email protected]. Behavioral fingerprints include consistent request intervals, absence of JavaScript execution, and reverse DNS hostnames such as crawler.wikio.com. Traffic analysis tools like useragentstring.com and botscout.com categorize it as a known aggregator bot under “Search Engine Crawlers”. Notably, the bot does not support cookies or session persistence.
📊 Data Usage
Collected data was primarily used to populate the Wikio news aggregation service, which categorized headlines from thousands of sources and powered the Wikio search engine for real‑time news searching. According to a 2010 company statement, aggregated content was not used for AI model training but solely for index building. After acquisition by eSky Group, data may have been repurposed for broader analytics, though no public evidence confirms AI training use. The bot’s focus remains on public news content, and it does not extract personal or non‑public information.
⚙️ Rate Limiting Policy
Although Wikio is a legitimate crawler with moderate request rates, it is rate‑limited by many webmasters due to occasional bursts during initial indexing of large sites. A threshold of 10 requests per second per IP is commonly recommended to prevent server overload while still allowing legitimate crawling. The bot’s adherence to Crawl‑Delay makes threshold‑based blocking effective and fair, ensuring that sites can protect their resources without entirely barring a reputable aggregator.
Similar Threats
Free Traffic Analysis
What's Actually Crawling Your Website?
Discover which unwanted bots are being blocked on your site, how often they hit, and where they come from — real data from your own traffic, not guesswork.
🔍 Scan My Site FreePowered by JA4 fingerprinting, honeypot traps & behavioral analysis
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.