omgili

Bot User-Agent: omgili

🤖 Overview

OmgiliBot is the web crawler operated by Omgili Ltd., an Israeli company founded in 2008 that runs Omgili.com, a specialized vertical search engine indexing only online forums, discussion boards, bulletin boards, and Q&A platforms. Its sole purpose is to aggregate publicly accessible user-generated content to enable cross-forum search functionality. The bot is mentioned in the Omgili about page and is documented at http://omgili.com/bot.html, confirming its legitimate, non-malicious intent.

🌐 Technical Behavior

OmgiliBot crawls by traversing links from forum sitemaps and link directories, typically fetching one page every 2–5 seconds to remain polite. It uses HTTP/1.1 with gzip compression and follows redirects. IP ranges are not publicly listed but historical observations show addresses originating from AWS EC2 (us-east-1 and eu-west-1) as well as Hetzner cloud. The bot does not execute JavaScript or parse AJAX content, focusing solely on static HTML. It respects the Crawl-Delay directive in robots.txt if set. Requests are made with a standard Accept header and sometimes a From header containing a contact email ([email protected]).

📋 robots.txt Compliance

Omgili explicitly states on its bot information page that OmgiliBot fully honors robots.txt Disallow rules and will also obey meta tags like noindex and nofollow. Site owners can completely block the bot by adding User-agent: OmgiliBot and Disallow: / to their robots.txt file. This compliance is verified by community reports from forum administrators.

🔍 Detection Indicators

The primary User-Agent string is OmgiliBot (case-sensitive). A longer variant appears as Mozilla/5.0 (compatible; OmgiliBot/1.0; +http://omgili.com/bot.html). No other custom headers are typically sent, but the bot may include a Via header when using proxies. Behavioral fingerprints include consistent inter-request delays and no referrer spoofing. The bot always identifies itself and never mimics other crawlers.

📊 Data Usage

Data collected by OmgiliBot is used exclusively to populate and update the Omgili search index, which allows users to search across millions of forum threads for discussions, product reviews, technical solutions, and opinions. The index is refreshed periodically but not in real-time. Omgili does not sell the data or use it for AI training; it is strictly a search engine service for public forum content. The company’s privacy policy confirms no data is retained beyond what is needed for indexing.

⚙️ Rate Limiting Policy

Web administrators are advised to rate-limit OmgiliBot if its default crawl rate strains server resources, though the bot is designed to self-throttle under load. A recommended threshold is 10 requests per second per IP range, with a 429 status code response to trigger back-off. This policy ensures fair resource usage without denying access to a legitimate indexing service.

⚠️

Your Site May Be Hemorrhaging Revenue to Bots

Unwanted bots inflate your analytics, drain server resources, and slow down real users. Check if your site is affected — completely free.

Check My Site for Free

Free to start  ·  Cancel anytime

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.