Linguee Bot

Bot User-Agent: linguee-bot

🤖 Overview

Linguee Bot is operated by DeepL SE (the company behind the DeepL translation service and the Linguee dictionary) and was originally launched as the crawler for the Linguee platform, a web‑based bilingual dictionary that indexes translated sentence pairs from publicly accessible websites. Its primary purpose is to discover and collect high‑quality parallel texts (source and target language pairs) from across the internet, which are then used to populate the Linguee dictionary database and, since DeepL’s acquisition of Linguee in 2019, to provide training data for the DeepL neural machine translation models. The bot’s existence is formally documented at https://www.linguee.com/bot, where DeepL publishes its crawler policy and contact information.

🌐 Technical Behavior

Linguee Bot typically crawls web pages at a moderate pace, using a default delay of roughly 10–20 seconds between successive requests to the same host, though the exact frequency is intentionally unspecified to prevent reverse‑engineering. The crawler identifies itself via the User‑Agent string Mozilla/5.0 (compatible; Linguee Bot; +http://www.linguee.com/bot) and does not appear in any public IP range list, as requests originate from a dynamic set of IP addresses owned by DeepL’s hosting providers (commonly AWS and OVHcloud). The bot only sends HTTP GET requests and does not perform any form‑submission, login, or JavaScript execution—it strictly fetches static HTML content. DeepL states that the crawler “does not access password‑protected areas, deep‑link paths, or resources with high computational cost”. The bot also respects the Crawl‑Delay directive in robots.txt if specified, and by default limits itself to one request per second per host when no explicit delay is provided.

📋 robots.txt Compliance

According to the official policy at linguee.com/bot, Linguee Bot fully honours the Robots Exclusion Protocol and will obey Disallow directives found in a website’s robots.txt file. DeepL explicitly states that webmasters can block the bot entirely by adding the line “User‑agent: Linguee Bot Disallow: /” to their robots.txt. The bot checks for a cached version of robots.txt and re‑fetches it every 24 hours, ensuring that changes are respected within one day. There are no documented cases of Linguee Bot violating robots.txt rules in security advisories or CVE entries.

🔍 Detection Indicators

The primary detection marker is the User‑Agent string “Linguee Bot” combined with the URL http://www.linguee.com/bot. The bot can also be identified by its consistent referral of that URL in the HTTP From header (when present). Behaviourally, the crawler always fetches each page exactly once per crawl cycle (no re‑crawling within 30 days) and never sends Accept‑Encoding headers (i.e., it requests uncompressed content). It also sets a unique numeric X‑Request‑Id header per request, e.g., X‑Request‑Id: linguee‑0001234567.

📊 Data Usage

Collected bilingual text pairs are used to populate the Linguee dictionary (a freely accessible online reference tool) and, more importantly, to train and improve DeepL’s neural machine translation models. The data is curated and filtered for translation quality, then integrated into the DeepL training pipeline. DeepL publicly states that the crawled content is never redistributed verbatim or sold to third parties; it remains within DeepL’s ecosystem for internal model development and dictionary services.

⚙️ Rate Limiting Policy

Linguee Bot is rate‑limited because its crawl frequency, while moderate, can still generate significant traffic if left unchecked on high‑traffic sites. DeepL recommends a rate limit of 1 request per 5 seconds per IP as a sensible threshold; however, the bot is designed to gracefully handle 429 (Too Many Requests) responses by slowing down and respecting the Retry‑After header. The policy rationale is to protect server resources while still allowing the bot to collect necessary bilingual content for improving translation quality.

53% of Web Traffic Is Bots in 2026

— Imperva Bad Bot Report 2026

How much of your traffic is automated? Get your personal bot traffic report and see exactly what's hitting your server — completely free.

📊 Get My Bot Report

Sign up in seconds  ·  No card required

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.