naver

Bot User-Agent: naver

🤖 Overview

Naver operates the NaverBot (also known as Yeti) web crawler, which is the primary indexing agent for the Naver Search Engine (www.naver.com), South Korea’s dominant search portal. According to Naver’s official documentation (searchadvisor.naver.com), the crawler is used to discover and index web pages for its search results, news aggregation, and AI-powered services including the HyperCLOVA X large language model. The bot is fully legitimate and follows standard crawling protocols as part of Naver’s webmaster tools.

🌐 Technical Behavior

NaverBot employs a distributed crawling infrastructure with IP addresses published in the NetRange WHOIS database under ASN AS23576 (Naver Business Platform) and AS45932. The crawler typically operates from IP ranges such as 125.209.208.0/20 and 203.232.224.0/19, which are geolocated in South Korea. It sends HTTP/1.1 requests with a user-agent string and often includes the header From: [email protected]. Crawl frequency is aggressive but adjustable; Naver’s guidelines recommend setting a crawl delay via robots.txt (e.g., Crawl-delay: 10 seconds). The crawler supports both HTTP and HTTPS, respects Host headers, and identifies itself via reverse DNS lookups that resolve to *.yetilab.net or *.naver.com.

📋 robots.txt Compliance

NaverBot fully respects robots.txt directives as documented in Naver’s webmaster guidelines (searchadvisor.naver.com). It honors Disallow rules for specific paths and also obeys the Crawl-delay directive. Additionally, Naver provides a separate control panel where site owners can manage crawling permissions and set exclude patterns. There is no evidence of intentional bypassing; the bot is strictly compliant with the Robots Exclusion Protocol.

🔍 Detection Indicators

The primary user-agent string is Mozilla/5.0 (compatible; NaverBot/1.0; +http://help.naver.com/robots/) for its main crawler, and Mozilla/5.0 (compatible; Yeti/1.1; +http://help.naver.com/robots/) for the Yeti variant used for additional data collection. Behavioral fingerprints include requests originating from South Korean IP ranges, consistent timing patterns, and the presence of the From header. Logs may also show a low TTL DNS cache with hostnames ending in yetilab.net.

📊 Data Usage

Collected data feeds Naver’s search index, the Naver Knowledge Encyclopedia, and training datasets for HyperCLOVA X — Naver’s generative AI model announced in 2023. The crawler also supports Naver Webmaster Tools for SEO analytics, and the data is used to improve search relevance, answer retrieval, and localized Korean-language AI services. Naver states it does not collect personal data or content behind authentication.

⚙️ Rate Limiting Policy

Due to its aggressive default crawl speed, many web administrators rate-limit NaverBot using server-level throttling (e.g., via mod_evasive or NGINX rate limits) to prevent resource exhaustion. The policy rationale is that while the bot is legitimate, its high request rate (sometimes exceeding 1 request per second per IP) can degrade performance for other visitors, justifying threshold-based blocking after repeated bursts beyond a reasonable crawl delay.

Free Traffic Analysis

What's Actually Crawling Your Website?

Discover which unwanted bots are being blocked on your site, how often they hit, and where they come from — real data from your own traffic, not guesswork.

🔍 Scan My Site Free

Powered by JA4 fingerprinting, honeypot traps & behavioral analysis

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.