pagepeeker

Bot User-Agent: pagepeeker

🤖 Overview

PagePeeker is a legitimate web screenshot service operated by the company PagePeeker Inc., publicly launched in 2010 and based in the Netherlands. Its primary purpose is to capture full-page screenshots of websites on demand, providing a reliable visual preview for use in link sharing, social media cards, bookmarking services, content management systems, and analytics dashboards. The service does not feed data into AI training models; instead, it generates static image snapshots for human or application consumption.

🌐 Technical Behavior

PagePeeker uses a headless browser engine—historically based on PhantomJS and later transitioned to Headless Chromium—to render and capture web pages. Requests are made with a configurable viewport size (default 1024×768) and a default timeout of 30 seconds. The crawler can follow redirects and execute JavaScript, meaning it may trigger dynamic content. IP ranges are not officially published, but reported blocks include 185.19.26.0/24 and 46.4.64.0/22 (Hetzner data centers in Germany). Request frequency depends on user subscription: free accounts are limited to 50 screenshots per hour; paid plans allow higher concurrency. PagePeeker sends standard HTTP GET requests and expects a 200 OK response; it will not crawl multiple pages unless explicitly requested via its API.

📋 robots.txt Compliance

According to PagePeeker’s official documentation (pagepeeker.com/bot), the service does respect robots.txt Disallow directives. It checks the robots.txt of the target domain before taking a screenshot. If a page is disallowed, the service returns an error or placeholder image. The documentation explicitly states: “We honor robots.txt rules just like a search engine.” This behavior is verified by multiple webmaster reports on community forums.

🔍 Detection Indicators

The primary User-Agent string is Mozilla/5.0 (compatible; PagePeeker/2.0; +http://pagepeeker.com/bot). Older versions used PagePeeker Screenshot Service 1.0. Behavioral fingerprints include a high frequency of requests from the same IP within a short span (depending on plan), a minimal set of HTTP headers (no Accept-Encoding for gzip in some cases), and the absence of a Referer header. The service also identifies itself via the X-PagePeeker HTTP header set to 1 in some API responses. Log analysis can use these indicators to distinguish PagePeeker from other screenshot bots like Googlebot or Applebot.

📊 Data Usage

Collected screenshots are stored temporarily on PagePeeker’s servers for up to 24 hours (free tier) or longer for paid users, then deleted. The images are used solely for visual preview generation and are not stored indefinitely, analyzed, or used to train machine learning models. PagePeeker’s privacy policy (pagepeeker.com/privacy) states they do not log the content of captured pages beyond the image itself, and do not share screenshots with third parties. The service is purely a utility tool for website previews.

⚙️ Rate Limiting Policy

PagePeeker is rate-limited because its screenshot requests can consume server resources and bandwidth, especially when triggered frequently by automated scripts. A threshold-based blocking policy—such as limiting the number of requests per minute from a single IP—is recommended to prevent accidental overloading, while still allowing legitimate screenshot captures. The official documentation advises webmasters to use robots.txt to disallow paths that should not be rendered, complementing rate limits.

Free Bot Analysis

Is Your Site Under Bot Attack Right Now?

Find out exactly how much of your traffic is automated — and which bots are draining your bandwidth and skewing your analytics.

Run Free Bot Scan →

No credit card required  ·  Results in minutes

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.