lexxebot
lexxebot is a web crawler operated by Lexxe, a Hong Kong‑based alternative search engine that has been running since 2005. Lexxe differentiates itself through natural‑language processing and semantic search capabilities, and lexxebot is the primary agent that indexes publicly accessible web pages for the Lexxe search index. According to the official Lexxe website (lexxe.com), the crawler is designed to respect website owners’ preferences while collecting content for search‑result freshness.
lexxebot typically crawls with a moderate request rate, but it can become aggressive on sites with frequently updated content, sending up to several requests per second during peak indexing cycles. The crawler uses standard HTTP/1.1 GET requests and follows redirects (301, 302) as well as meta refresh tags. Its IP addresses originate primarily from Hong Kong and surrounding Asian regions, though Lexxe has not published a formal IP range list. The bot honors the Crawl‑Delay directive if set in robots.txt, but may ignore it when the delay is unreasonably high. It also respects noindex meta tags and nofollow link attributes.
Lexxe’s official documentation (available on lexxe.com/robots) states that lexxebot fully honors all standard robots.txt directives, including Disallow, Allow (with partial support for wildcards), and Crawl‑Delay. Third‑party tests (e.g., from the Robots Exclusion Protocol community) confirm that the bot does not access disallowed paths unless the directive is misconfigured. However, active monitoring by site administrators reports that lexxebot occasionally re‑crawls pages shortly after the robots.txt file is updated, suggesting a slight delay in policy propagation.
The primary identifying User‑Agent string is LexxeBot/1.0 (or simply lexxebot), often formatted as: "Mozilla/5.0 (compatible; LexxeBot/1.0; +http://www.lexxe.com/)". Some versions omit the Mozilla prefix. Lexxe does not send custom HTTP headers like X‑Crawler‑Identity, but the bot’s unique request pattern — sending the same User‑Agent consistently and rarely including an Referer header — aids detection. Log entries from the Lexxe crawl usually show a distinct IP range around 202.125.*.* (Hong Kong), though the range shifts periodically.
The data collected by lexxebot is used exclusively to populate and update the Lexxe search engine index. Lexxe’s privacy policy states that cached copies of web pages may be stored temporarily, but full content is not retained beyond the indexing process. Unlike many AI‑training bots, Lexxe does not currently share crawled data with third parties or use it for language‑model training; the data is solely for search‑result ranking and snippet generation.
lexxebot is rate‑limited because its moderate crawl rate can still overwhelm small websites with limited server resources, especially when multiple threads are active simultaneously. A threshold‑based blocking policy (e.g., restricting to 1 request per 10 seconds) is recommended to ensure the site remains responsive to human visitors while still allowing legitimate indexing.
Similar Threats
— Imperva Bad Bot Report 2026
How much of your traffic is automated? Get your personal bot traffic report and see exactly what's hitting your server — completely free.
📊 Get My Bot ReportSign up in seconds · No card required
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.
Stay up to date with the latest from Boteraser.
We use cookies to improve your experience on our site. By using our site, you consent to cookies.
Manage your cookie preferences below:
Essential cookies enable basic functions and are necessary for the proper function of the website.
CloudFlare provides web performance and security solutions, enhancing site speed and protecting against threats.
Service URL: developers.cloudflare.com (opens in a new window)
These cookies are needed for adding comments on this website.
These cookies are used for managing login functionality on this website.
Statistics cookies collect information anonymously. This information helps us understand how visitors use our website.
Google Analytics is a powerful tool that tracks and analyzes website traffic for informed marketing decisions.
Service URL: policies.google.com (opens in a new window)
You can find more information in our Cookie Policy and Privacy Policy.