Skip to main content

Boteraser | Website and Server Security Solutions

omgilibot

Bot User-Agent: omgilibot

🤖 Overview

omgilibot is a web crawler operated by Omgili Ltd. (now part of BoardReader), a search engine specializing in indexing content from online forums, discussion boards, and community platforms. Its primary purpose is to collect publicly accessible forum threads and posts to power Omgili’s search results, which allow users to find relevant discussions across thousands of independent forums. The bot was first documented in the early 2000s and remains active as of 2024, with official documentation available at http://www.omgili.com/bot.html (archived via Internet Archive).

🌐 Technical Behavior

omgilibot typically crawls at a moderate frequency, sending requests in bursts to avoid overwhelming smaller forum servers. It follows HTML links and respects noindex meta tags, and it often fetches pages sequentially using HTTP/1.1 with a standard GET method. The bot’s IP ranges are not publicly documented by Omgili, but historical data shows it originates from a small set of IP addresses registered to Omgili Ltd. in Israel. It does not use JavaScript rendering and only parses static HTML content, making it lightweight for most web servers. The crawler supports compression (Accept-Encoding: gzip) and includes a From header with an administrative email address for contact.

📋 robots.txt Compliance

According to Omgili’s official bot page, omgilibot fully honors robots.txt directives and will cease crawling any URL or directory that includes a Disallow rule. This is confirmed by forum administrators who have tested the bot’s behavior; it also respects Crawl-Delay directives when specified. However, some site owners have noted that the bot may ignore Disallow when the robots.txt file is temporarily unreachable, though this is not documented as intentional.

🔍 Detection Indicators

The primary User-Agent string is omgilibot/1.0 (+http://www.omgili.com/bot.html). Behavioral fingerprints include requests to forum-specific URL patterns (e.g., /showthread.php, /viewtopic.php) and a low average of 1–2 requests per second. The bot also sends a Via header indicating the proxy used. No secondary User-Agent aliases are known, making detection straightforward using log analysis.

📊 Data Usage

Collected data—forum thread titles, post bodies, author names, and timestamps—is stored and indexed by Omgili’s search engine to provide a cross-forum search service. The data is not used for AI training or large language models; it is purely for retrieval and ranking of forum discussions. Omgili does not sell or redistribute raw content, but aggregated search snippets are displayed publicly. The index is updated regularly to reflect new posts.

⚙️ Rate Limiting Policy

omgilibot is rate-limited by many forum administrators because its crawling, while respectful, can still generate a significant load on high-traffic discussion boards. The policy rationale is to prevent resource exhaustion while allowing the bot to maintain a fresh index—thresholds of 5–10 requests per second are commonly applied, with automatic blocking above that level to protect server stability.

Free Traffic Analysis

What's Actually Crawling Your Website?

Discover which unwanted bots are being blocked on your site, how often they hit, and where they come from — real data from your own traffic, not guesswork.

🔍 Scan My Site Free

Powered by JA4 fingerprinting, honeypot traps & behavioral analysis

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.