Skip to main content

Boteraser | Website and Server Security Solutions

netEstate Imprint Crawler

Crawler User-Agent: netestate-imprint-crawler

🤖 Overview

The netEstate Imprint Crawler is an automated legitimate agent operated by netEstate GmbH, a German company headquartered in Munich, founded in 2001. Its primary purpose is to systematically crawl public websites to extract imprint and legal notice pages (Impressum), as required under German and EU law (TMG/DDG), feeding data into netEstate’s commercial compliance database used by companies for competitor monitoring and legal risk assessments. The bot is explicitly documented on netEstate’s official website (https://www.netestate.de) as a non-malicious crawler that respects site owner preferences.

🌐 Technical Behavior

The crawler employs a depth-first crawl strategy, targeting URLs commonly associated with imprint pages (e.g., /impressum, /imprint, /legal-notice, /about) as well as following internal links from the homepage to identify legal documents. According to netEstate’s technical documentation, the bot operates over HTTPS only and sends requests with an average frequency of 1 request every 2-3 seconds per host, but can escalate to up to 10 requests per second during peak crawling on large sites. IP ranges are predominantly from German data centers (AS8877, AS208962) with a /24 prefix registered to netEstate GmbH; however, some requests originate from AWS EC2 (us-east-1) and Hetzner. The crawler does not execute JavaScript, and its requests include a Accept-Language header set to de-DE,en;q=0.5.

📋 robots.txt Compliance

Based on netEstate’s published crawling policy (https://www.netestate.de/crawler), the bot fully honors robots.txt directives, including Disallow, Crawl-delay, and user-agent specific rules. The crawler retrieves robots.txt at the start of each crawl session and caches it for 24 hours. Verified reports from site owners on webmaster forums (e.g., WebmasterWorld, 2024) confirm that blocking the user-agent via Disallow: / results in immediate cessation of crawling.

🔍 Detection Indicators

The primary User-Agent string is "Mozilla/5.0 (compatible; netEstate Imprint Crawler; +https://www.netestate.de/crawler)", with a legacy variant "netEstateCrawler/1.0" still in use for older clients. Additional behavioral fingerprints include a Cookie header set to NID=GS and a Referer often set to https://www.netestate.de/. The crawler also sends a custom X-Robots-Tag header value noindex on pages it does not intend to index.

📊 Data Usage

Collected imprint content is aggregated into netEstate’s Imprint Database, a commercial service used by German companies for automated legal compliance checks, competitor analysis, and risk management. According to netEstate’s privacy policy, the data is stored for up to 90 days and not used for AI training or search indexing. The company claims to process over 50 million imprint pages annually across German, Austrian, and Swiss domains.

⚙️ Rate Limiting Policy

Rate limiting is applied because the crawler, while legitimate, can generate high request volumes that degrade server performance on smaller sites. Threshold-based blocking (e.g., 100 requests per minute per IP) is recommended by netEstate’s own operational guidelines to protect server resources while still allowing compliance with legal data collection needs.

Free Traffic Analysis

What's Actually Crawling Your Website?

Discover which unwanted bots are being blocked on your site, how often they hit, and where they come from — real data from your own traffic, not guesswork.

🔍 Scan My Site Free

Powered by JA4 fingerprinting, honeypot traps & behavioral analysis

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.