Piepmatz

Bot User-Agent: piepmatz

🤖 Overview

The Piepmatz crawler is operated by Piepmatz GmbH, a German company founded in 2021 that runs a privacy-focused meta‑search engine of the same name. Its primary purpose is to build an independent web index that respects user anonymity and does not store query logs. The bot collects publicly available web content to fuel the Piepmatz search engine, which aggregates results from multiple back‑end sources without retaining personal data.

🌐 Technical Behavior

The crawler uses a distributed architecture based on Apache Nutch and Solr, as disclosed in the Piepmatz technical blog (blog.piepmatz.com/crawling). It sends HTTP/1.1 requests with a default delay of 5 seconds between consecutive hits to the same domain, but can be tuned via the Crawl‑Delay directive in robots.txt. The bot rotates through a pool of IP addresses owned by Hetzner Online GmbH and Hetzner Cloud (German data centers), identifiable by reverse DNS names like crawler‑pz‑*.your‑server.de. Requests are always made over IPv4; no IPv6 support has been observed. The crawler fetches only text/html and text/plain content, ignoring images, PDFs, and other binary files, as stated in its source code repository (github.com/piepmatz/crawler).

📋 robots.txt Compliance

Piepmatz fully respects the Robots Exclusion Protocol, including both Disallow and Allow directives. The official documentation on piepmatz.com/robots.txt-policy confirms that any path blocked via Disallow is never visited. Additionally, the bot honors the Crawl‑Delay directive if set; otherwise it falls back to its internal 5‑second interval.

🔍 Detection Indicators

The primary User‑Agent string is Mozilla/5.0 (compatible; Piepmatz/1.0; +https://piepmatz.com/crawler), occasionally supplemented with From: [email protected] in HTTP headers. A secondary variant Piepmatz‑Bot/1.0 has been observed on older deployments. The bot always includes a X‑Crawler‑Version: Piepmatz header and a valid User‑Agent token. No obfuscation or spoofing is documented.

📊 Data Usage

Collected data feeds the Piepmatz search index, which powers anonymous web searches. The company explicitly states that it does not use crawled content for AI training, advertising profiling, or third‑party data sales. All indexing is performed on servers located in Germany and subject to the GDPR. Per the privacy policy (piepmatz.com/privacy), raw crawl data is discarded after 30 days and only anonymized metadata is retained.

⚙️ Rate Limiting Policy

While Piepmatz operates legitimately, it can still consume significant bandwidth when indexing large sites because it performs deep, recursive crawling. Rate limiting is therefore recommended to ensure fair resource allocation and to prevent the bot from overwhelming server capacity, especially during peak hours or on shared hosting environments.

53% of Web Traffic Is Bots in 2026

— Imperva Bad Bot Report 2026

How much of your traffic is automated? Get your personal bot traffic report and see exactly what's hitting your server — completely free.

📊 Get My Bot Report

Sign up in seconds  ·  No card required

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.