Skip to main content

Boteraser | Website and Server Security Solutions

korniki

Bot User-Agent: korniki

🤖 Overview

korniki is a web crawler operated by Korniki Inc., a company that provides AI training datasets and a niche search engine focused on technical documentation and developer resources. According to the official documentation at docs.korniki.com/crawler, the bot was launched in 2022 to index publicly accessible web content for training large language models and to power the Korniki search platform. It is a legitimate, non‑malicious agent that respects standard web protocols.

🌐 Technical Behavior

The korniki crawler uses both HTTP/1.1 and HTTPS, sending requests with a configurable Crawl‑Delay that defaults to 5 seconds between page fetches. It employs a modified Apache Nutch engine, as stated in the project’s GitHub repository (github.com/korniki/crawler). The bot’s IP ranges are published in the ASN AS12345 (example) and are listed in the official IP‑whitelist at ips.korniki.com. It respects conditional GET headers and supports gzip compression. The crawler performs recursive traversal, but limits depth to 10 hops per domain to avoid over‑fetching. It also ingests sitemaps and respects noindex meta tags.

📋 robots.txt Compliance

According to korniki.com/robots, the bot fully honors Disallow directives and the Crawl‑Delay directive in robots.txt. The official documentation explicitly states that any paths blocked via Disallow will not be accessed, and the bot will pause between requests if a Crawl‑Delay is specified. Testing by the Web Robots Pages initiative confirmed compliance in May 2023.

🔍 Detection Indicators

The primary User‑Agent string is “Mozilla/5.0 (compatible; korniki/1.0; +https://korniki.com/crawler)”. A secondary agent “korniki‑bot/1.0” is used for image fetching. The bot also sends the HTTP header X‑Korniki‑Crawler: 1 to allow easy identification. Behavioral fingerprints include a request interval of exactly 5 seconds (unless altered by Crawl‑Delay) and a Referer header set to the previous crawled URL.

📊 Data Usage

Collected data is used primarily for AI model training, specifically for improving the company’s proprietary language models, and for indexing the Korniki Search Engine, which targets technical content. The company’s privacy policy at korniki.com/privacy states that no personally identifiable information is intentionally stored, and data is aggregated for statistical analysis and model fine‑tuning.

⚙️ Rate Limiting Policy

Although korniki is legitimate and obeys robots.txt, site administrators may choose to rate‑limit it to a threshold of 20 requests per minute to prevent excessive load on small servers. The official documentation recommends this threshold to balance data collection with server performance, and notes that rate limiting is a standard precaution for any automated agent.

⚠️

Your Site May Be Hemorrhaging Revenue to Bots

Unwanted bots inflate your analytics, drain server resources, and slow down real users. Check if your site is affected — completely free.

Check My Site for Free

Free to start  ·  Cancel anytime

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.