ozelot
Bot User-Agent:ozelot
🤖 Overview
Ozelot is the web crawler operated by Mojeek, an independent, UK-based search engine that prioritizes user privacy and does not employ tracking or personalization. First deployed in the early 2010s, Ozelot’s primary purpose is to discover and index publicly accessible web pages for Mojeek’s search results, offering an alternative index that is not reliant on major search engines like Google or Bing. The crawler feeds data exclusively into Mojeek’s proprietary search index, which is built from scratch and maintained independently of third-party sources.
🌐 Technical Behavior
Ozelot performs regular, periodic crawls using a breadth-first traversal strategy. It sends requests with the standard HTTP/1.1 protocol and respects the Accept-Encoding header for gzip and deflate compression. The crawler’s IP addresses are assigned from netblocks owned by Mojeek, typically originating from UK-based data centers such as those operated by OVHcloud and Hetzner. According to Mojeek’s official documentation, the crawl frequency is moderate — pages are re-visited based on observed change frequency, with a default crawl-delay of 1 second between requests to the same host. Ozelot follows rel="canonical" links and respects nofollow directives. It does not attempt to access hidden directories or password-protected areas, and it parses sitemaps if referenced in robots.txt or submitted via a Sitemap ping.
📋 robots.txt Compliance
Mojeek explicitly states that Ozelot fully obeys the Robots Exclusion Protocol as published on their official website. The crawler reads and honors both Disallow and Crawl-delay directives in robots.txt files. Mojeek also provides a dedicated page (mojeek.com/about/crawler) where webmasters can find instructions for blocking or customizing access, confirming that Ozelot does not deliberately circumvent any exclusion rules.
🔍 Detection Indicators
The primary User-Agent string used by Ozelot is "Mozilla/5.0 (compatible; Ozelot/1.0)". In some configurations, the string may appear as "MojeekBot/1.0" or include variations like "Mojeek/1.0". The bot also identifies itself via a User-Agent: Ozelot header in some requests. Behavioral fingerprints include a consistent request pattern with a 1-second minimum interval per host, and the use of HTTP/1.1 without pipelining. No additional identifying headers are typically sent beyond the standard Host and Accept fields.
📊 Data Usage
All data collected by Ozelot is used exclusively for building and updating Mojeek’s search index. Mojeek does not use the crawled content for training AI models, machine learning, or any secondary purposes such as advertising analysis. The index is publicly queryable through Mojeek’s search engine and is periodically refreshed to reflect changes in the web. Mojeek’s privacy policy explicitly states that no personal data is harvested from crawled pages beyond what is necessary for indexing.
⚙️ Rate Limiting Policy
Rate limiting is applied because Ozelot’s crawl can be aggressive if left unchecked, potentially degrading server performance for smaller sites. Mojeek recommends that webmasters set a Crawl-delay directive in robots.txt or use Disallow to block sections they do not wish to be indexed. The policy rationale is to balance indexing completeness with server load, with threshold-based blocking used only when a bot exceeds 5 requests per second or receives HTTP 429 (Too Many Requests) responses.
Similar Threats
Free Traffic Analysis
What's Actually Crawling Your Website?
Discover which unwanted bots are being blocked on your site, how often they hit, and where they come from — real data from your own traffic, not guesswork.
🔍 Scan My Site FreePowered by JA4 fingerprinting, honeypot traps & behavioral analysis
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.