weblexbot

Bot User-Agent: weblexbot

🤖 Overview

weblexbot is a web crawler operated by WebLex GmbH, a German legal information provider based in Switzerland (weblex.de). Its primary purpose is to systematically collect and index publicly available legal documents, court rulings, statutes, and academic legal publications to feed into the Weblex legal research platform. The bot was first documented in public robots.txt files around 2018 and is explicitly listed on the weblex.ch website as the official crawler for their service.

🌐 Technical Behavior

weblexbot initiates HTTP/1.1 GET requests with a default crawl frequency of approximately one request every 10 to 15 seconds per domain, though the rate may vary depending on server response times. Official documentation from weblex.ch states that the crawler uses the Accept-Encoding: gzip header to reduce bandwidth usage and supports ETag and If-Modified-Since conditional requests. The IP ranges are not publicly aggregated but are geolocated primarily in Switzerland and Germany, with some hosts resolving to the Swiss hosting provider Exoscale. The bot only crawls on port 80 (HTTP) and port 443 (HTTPS) and does not follow redirects that lead to non‑HTTP protocols.

📋 robots.txt Compliance

Based on the official Weblex crawler policy page at weblex.ch/robots.txt and independent testing reported on user forums, weblexbot fully respects Disallow directives in robots.txt. It also obeys Crawl-Delay directives if present, pausing for the specified number of seconds between requests. There are no documented cases of the bot ignoring robots.txt rules, and the operator has published a contact email ([email protected]) for site owners to request adjustments if needed.

🔍 Detection Indicators

The definitive User-Agent string for weblexbot is "weblexbot/1.0" (occasionally with a version suffix like "1.0.1"). The bot does not disguise itself and always includes a From header with the email "[email protected]" for administrative contact. Additionally, it sets a User‑Agent that contains no other crawler names, making it easily distinguishable in server logs.

📊 Data Usage

All data collected by weblexbot is used exclusively for the Weblex legal research platform, which provides full‑text search, citation analysis, and case‑law connections for legal professionals. The crawler does not train AI models or generate derivative works outside the Weblex service. Retrieved documents are stored in a structured database and are updated whenever the source document changes, as indicated by HTTP last‑modified timestamps.

⚙️ Rate Limiting Policy

weblexbot is rate‑limited because, while it is legitimate and respects robots.txt, its continuous crawling can consume significant server resources on small sites hosting legal content. A threshold‑based blocking mechanism (e.g., limiting to 30 requests per minute per IP) is a reasonable defense to prevent unintended load without denying the bot access altogether.

🛡️

Stop Bots. Save Bandwidth. Protect Revenue.

Boteraser automatically detects and blocks unwanted bots — protecting your site from scrapers, DDoS bursts, and credential stuffing attacks without slowing down real visitors.

✅ Start Free Protection

Setup takes under a minute  ·  Free trial available

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.