twengabot

Bot User-Agent: twengabot

🤖 Overview

twengabot is an automated web crawler operated by Twenga, a Paris-based price comparison and e-commerce aggregator founded in 2006. The bot systematically indexes product listings, prices, availability, and merchant details from thousands of online retail websites to feed Twenga’s shopping search engine, which serves users in France, Germany, Spain, Italy, the UK, and several other European countries. According to Twenga’s official documentation and the robots.txt guidelines published on their site (twenga.com/robots.txt), the bot is designed exclusively for the purpose of collecting publicly available e-commerce data and does not attempt to access protected or non-public content.

🌐 Technical Behavior

twengabot performs HTTP GET requests with a configurable crawl delay, typically ranging from 1 to 10 seconds between requests, as documented in Twenga’s webmaster resources. The crawler follows standard HTTP/1.1 and HTTPS protocols, employs gzip compression for efficiency, and respects Last-Modified and ETag headers to minimize bandwidth usage. Its IP ranges are predominantly assigned to Twenga’s own ASN (AS203876) and include IPv4 blocks from French and German hosting providers such as OVH and Hetzner; a sample of known IP addresses includes 185.15.58.x and 91.121.132.x based on community logs. The bot does not use JavaScript rendering and only fetches static HTML, JSON-LD structured data, and XML feeds (e.g., Google Shopping format) when explicitly linked. Requests are made with a randomized User-Agent variation that includes the string “twengabot” followed by a version number, and the bot identifies itself via the User-Agent header. Twenga maintains a public list of their crawler’s IP ranges on their official webmaster page (twenga.com/webmaster/crawler).

📋 robots.txt Compliance

twengabot is documented to fully honor robots.txt Disallow directives, including both specific path exclusions and wildcard patterns. Twenga explicitly instructs webmasters that the bot will cease crawling any URL that returns a 403, 404, or 410 status code, and it will cache the robots.txt file for up to 24 hours before rechecking. This behavior is confirmed by multiple independent webmaster forums and server logs that show twengabot never accesses protected areas like /cart, /admin, or /my-account when those paths are disallowed.

🔍 Detection Indicators

The primary identifying User-Agent string is “twengabot/1.0” (or variants like “twengabot/2.1”), occasionally accompanied by a comment such as “+https://www.twenga.com/bot.html”. Some requests carry a From header containing “[email protected]” for direct contact. Behavioral fingerprints include a consistent request interval (default 2 seconds), absence of other browser-like headers (e.g., Accept-Language is usually missing), and a narrow range of HTTP methods (only GET). Server operators can also detect the bot by its reverse DNS hostnames, which typically resolve to “crawler-[id].twenga.com”.

📊 Data Usage

Collected data — including product titles, prices, images, descriptions, stock status, and merchant names — is ingested into Twenga’s centralized product database to power real-time price comparison queries on their website and mobile apps. The data is also used for analytics reports provided to registered merchants, such as price positioning and availability trends. Twenga explicitly states that they do not sell raw crawled data to third parties; instead, the aggregated comparison service is monetized through affiliate commissions and advertising.

⚙️ Rate Limiting Policy

twengabot is rate-limited because its aggressive, automated crawling of product pages can consume significant server resources at high concurrency; standard practice involves throttling to 1 request per 5–10 seconds per IP. This policy is justified by Twenga’s own recommendation to webmasters who wish to reduce load, and it aligns with the principle that legitimate crawlers should not degrade site performance for human users while still fulfilling their indexing purpose.

🛡️

Stop Bots. Save Bandwidth. Protect Revenue.

Boteraser automatically detects and blocks unwanted bots — protecting your site from scrapers, DDoS bursts, and credential stuffing attacks without slowing down real visitors.

✅ Start Free Protection

Setup takes under a minute  ·  Free trial available

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.