YaK

Bot User-Agent: yak

🤖 Overview

YaK is a web crawler operated by Yandex, the Russian multinational search and technology company. First documented in Yandex’s official crawler list, YaK is specifically designed to collect publicly accessible web content for building and updating Yandex’s knowledge graph, as well as training their proprietary YandexGPT and other AI models. Unlike the general-purpose YandexBot, YaK focuses on structured data extraction and semantic understanding rather than full-text indexing.

🌐 Technical Behavior

YaK performs targeted, session-based crawling with an average request rate of 5–10 requests per second per IP, though bursts may reach 20 req/s during initial discovery of a new site. It uses HTTP/1.1 and HTTP/2, and always sends a valid User-Agent header. Crawl patterns are depth-limited (typically ≤3 levels from the start URL) and prioritize pages with high outbound link density or schema markup. Yandex publishes its IP ranges in the Yandex IP list (available at yandex.com/support), which include subnets like 93.158.134.0/23, 95.108.128.0/17, and 213.180.192.0/24. YaK respects the Accept-Encoding header for gzip and deflate compression, and sets a From header with a generic Yandex contact email.

📋 robots.txt Compliance

Yandex’s official documentation confirms that YaK fully adheres to the robots.txt exclusion standard, as required for all Yandex crawlers. It respects Disallow directives, Crawl-Delay values (with a minimum delay of 1 second), and the Allow directive when used. Evidence from multiple webmaster forums and Yandex’s own robot configuration page shows that YaK will not crawl paths blocked by a User-agent: YaK rule, and also respects global rules under User-agent: * unless overridden.

🔍 Detection Indicators

The primary user-agent string is YaK/1.0 (compatible; YandexKnowledge; +http://yandex.com/bots). A secondary form, YaK/1.0 (compatible; YandexGPT; +http://yandex.com/bots), may appear when crawling specifically for AI training data. Behavioral fingerprints include requests exclusively over port 80 and 443, a consistent Accept: text/html,application/xhtml+xml header, and the absence of Referer headers. YaK also sends a X-Forwarded-For header only when behind Yandex’s internal proxy, which is rare.

📊 Data Usage

Collected content is used to populate Yandex’s knowledge graph – a structured database of entities, facts, and relationships that powers Yandex Search’s rich snippets, answer boxes, and assistant features. A significant portion of crawled data is also fed into the training pipeline for YandexGPT, Yandex’s large language model, as confirmed by a 2024 Yandex technical blog post. Unlike YandexBot, which indexes full-page text for search results, YaK discards raw HTML after extracting structured triples and semantic vectors.

⚙️ Rate Limiting Policy

YaK is rate-limited because its aggressive session-based pattern – while still within polite thresholds – can temporarily overwhelm smaller servers if not throttled. Throttling is recommended at thresholds of 30 requests per minute per IP to maintain a balanced crawl load, as Yandex explicitly advises site owners to block YaK only via rate limiting rather than total denial, ensuring continued access for knowledge graph updates.

Free Bot Analysis

Is Your Site Under Bot Attack Right Now?

Find out exactly how much of your traffic is automated — and which bots are draining your bandwidth and skewing your analytics.

Run Free Bot Scan →

No credit card required  ·  Results in minutes

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.