empas_robot
Bot User-Agent:empas-robot
🤖 Overview
empas_robot is a web crawler operated by NHN Corporation (formerly Empas, a South Korean internet portal and search engine). First documented in the early 2000s, its primary purpose is to index Korean-language web content for the Empas search engine (empas.com), which later integrated into NHN’s broader search ecosystem. The bot collects publicly accessible pages to build the search index that powers query results for users in South Korea.
🌐 Technical Behavior
empas_robot follows standard HTTP/1.1 protocols with a User-Agent string of empas_robot (case‑sensitive). Crawl frequency is moderate, typically issuing requests every few seconds per domain, with a default crawl delay of 5 seconds when not explicitly set via Crawl‑Delay directives. IP ranges are assigned from NHN’s Korean netblocks (e.g., 210.118.x.x and 203.241.x.x). The bot respects Last‑Modified headers and ETag caching to minimise redundant downloads. It uses HTTP/1.1 Keep‑Alive connections and sends a Referer header pointing to http://www.empas.com on each request.
📋 robots.txt Compliance
empas_robot fully honours Disallow directives in robots.txt as documented by NHN’s official search guidelines (available on the Empas help page). It also respects the Crawl‑Delay directive, allowing webmasters to throttle its request rate. There are no known cases of the bot ignoring robot exclusion rules.
🔍 Detection Indicators
The primary detection string is the exact User‑Agent empas_robot. Some variants also include empas/1.0 or empas_robot/1.0 (rare). Behavioural fingerprints include a lack of Accept‑Encoding headers (no gzip support) and a fixed request interval of 5 seconds unless overridden. The bot also sends a From header containing the email address [email protected] on some requests.
📊 Data Usage
Collected content is used exclusively for web indexing and search result ranking for the Empas search engine. No data is repurposed for AI training, advertising profiling, or third‑party analytics. NHN’s privacy policy explicitly states that crawled data is not shared externally and is retained only as long as needed to maintain the search index.
⚙️ Rate Limiting Policy
empas_robot is rate‑limited because its crawl frequency, while moderate, can still overwhelm smaller servers if left unchecked. A threshold‑based block (e.g., >10 requests per second from its IP range) is recommended to protect site performance while allowing genuine indexing.
Similar Threats
🛡️
Stop Bots. Save Bandwidth. Protect Revenue.
Boteraser automatically detects and blocks unwanted bots — protecting your site from scrapers, DDoS bursts, and credential stuffing attacks without slowing down real visitors.
✅ Start Free ProtectionSetup takes under a minute · Free trial available
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.