gethtmlcontents

Bot User-Agent: gethtmlcontents

🤖 Overview

gethtmlcontents is a legitimate web crawler operated by the SaaS platform GetHTMLContents.com, a service designed to fetch and render complete HTML documents from user-supplied URLs for use in website previews, SEO audits, and content analysis tools. First publicly documented in official documentation on the vendor’s site (gethtmlcontents.com/robots.txt), the bot’s primary purpose is to provide up-to-date, rendered HTML snapshots for paying customers, enabling features like live site previews and broken-link checking.

🌐 Technical Behavior

According to the official documentation and observed traffic patterns reported in web server logs, gethtmlcontents uses a headless Chromium engine to execute JavaScript and render dynamic content, making it one of the few bots that fully processes modern single-page applications. It typically issues requests at a rate of 5–10 requests per second per IP, with bursts during peak customer usage. The bot primarily connects from IP ranges belonging to Amazon Web Services (AWS) and DigitalOcean, as confirmed by reverse DNS lookups. It uses HTTP/1.1 and HTTP/2 protocols and fetches all resources (CSS, JS, images) to recreate a complete page view. A notable technical detail is its use of a custom X-GetHTMLContents-Request header to identify each customer’s session (source: gethtmlcontents.com/docs/bot-behavior).

📋 robots.txt Compliance

The vendor explicitly states on its robots.txt page (gethtmlcontents.com/robots.txt) that the bot fully honors Disallow directives and respects Crawl-Delay directives when present. Evidence from multiple website administrators’ forum posts (e.g., WebmasterWorld thread #2023‑12) confirms that the bot stops crawling paths listed in robots.txt within 24 hours of a directive change. However, it does not respect noindex meta tags because it only fetches HTML, not index the content.

🔍 Detection Indicators

The primary User-Agent string is Mozilla/5.0 (compatible; gethtmlcontents/1.0; +https://gethtmlcontents.com/bot). Secondary identifiers include a custom X-Bot-Type: gethtmlcontents HTTP header and a consistent Accept-Language: en-US header. Behavioral fingerprints include extremely fast page loads (under 1 second) and the simultaneous fetching of all subresources from the same origin. Official documentation also lists alternative UA strings for different service tiers, such as gethtmlcontents-preview/2.0 (source: gethtmlcontents.com/user-agent).

📊 Data Usage

Collected HTML is not used for AI training, search indexing, or data mining. Instead, it is temporarily cached (for up to 48 hours, per their privacy policy) and served back to the customer who requested it via an API endpoint for use in visual regression testing, content comparison, and dynamic preview generation. The bot does not store any personal information or analyze text beyond what is necessary to render the page.

⚙️ Rate Limiting Policy

Although gethtmlcontents is a legitimate, non‑malicious agent, it can generate high volumes of traffic during peak customer demand, potentially overwhelming small sites. Therefore, rate‑limiting to 5 requests per second per IP is recommended by the vendor itself in its best‑practices guide (gethtmlcontents.com/rate-limiting), and threshold‑based blocking is a sensible precaution to preserve server resources without violating the bot’s intended purpose.

🛡️

Stop Bots. Save Bandwidth. Protect Revenue.

Boteraser automatically detects and blocks unwanted bots — protecting your site from scrapers, DDoS bursts, and credential stuffing attacks without slowing down real visitors.

✅ Start Free Protection

Setup takes under a minute  ·  Free trial available

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.