Skip to main content

Boteraser | Website and Server Security Solutions

keybot translation-search-machine

Search Engine User-Agent: keybot-translation-search-machine

🤖 Overview

keybot translation-search-machine is a legitimate web crawler operated by Keybot LLC (keybot.com), a company specializing in multilingual search and machine translation services. First documented in public User‑Agent lists in early 2022, the bot’s primary purpose is to collect web content for building a cross‑language search index that enables users to query documents in one language and retrieve relevant results in another, leveraging neural machine translation models trained on crawled data.

🌐 Technical Behavior

The bot performs HTTP/1.1 and HTTP/2 requests with a default crawl rate of approximately 10 requests per second per domain, though it may burst to 20 requests for brief periods. It respects ETag and Last‑Modified headers for incremental crawling. IP ranges are allocated from Keybot’s own ASN (AS208722), which includes the block 185.199.108.0/24 and 2a05:dfc0::/32 (IPv6). The crawler uses a custom Transaction‑ID header containing a UUID for request tracking. It does not support cookies or JavaScript rendering; it only fetches static HTML and plain‑text content. Crawl depth is limited to three levels from the entry point unless a sitemap is provided.

📋 robots.txt Compliance

According to the official documentation at keybot.com/robots-policy, the bot fully honors Disallow directives from robots.txt and also respects Crawl‑Delay instructions. However, it ignores Allow directives that conflict with a Disallow (it takes the most restrictive interpretation). The bot does not crawl URLs that return a 4xx or 5xx status code more than twice in a rolling 24‑hour window.

🔍 Detection Indicators

The primary User‑Agent string is keybot-translation-search-machine (without spaces). A secondary UA is Mozilla/5.0 (compatible; Keybot/2.0; +https://keybot.com/bot). Behavioral fingerprints include a mandatory X‑Keybot‑Request header set to 1 in all requests, and the bot always includes a From header with the email [email protected]. Requests originate from the official ASN and never spoof User‑Agent strings.

📊 Data Usage

Collected content is processed into a bilingual corpus for training domain‑specific translation models, as described in Keybot’s 2023 white paper “Cross‑Language Search via Neural Crawling”. The index supports fuzzy matching and transliteration for languages without Latin scripts. Keybot also offers a paid API (keybot.com/translate) that uses this data to provide contextual translations for web applications.

⚙️ Rate Limiting Policy

Rate‑limiting is recommended because the bot’s burst behavior can briefly consume server resources; a threshold of 30 requests per second with a 60‑second sliding window provides sufficient headroom while preventing overuse of server connections. This aligns with the bot’s own documented “polite” crawl rate of 10 rps and ensures protection for shared hosting environments.

🛡️

Stop Bots. Save Bandwidth. Protect Revenue.

Boteraser automatically detects and blocks unwanted bots — protecting your site from scrapers, DDoS bursts, and credential stuffing attacks without slowing down real visitors.

✅ Start Free Protection

Setup takes under a minute  ·  Free trial available

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.