Sitebeam

Bot User-Agent: sitebeam

🤖 Overview

Sitebeam is a legitimate web crawler operated by Sitebeam Ltd., a UK-based company (sitebeam.co), designed to perform automated SEO audits and website health checks for subscription customers. Its primary purpose is to crawl user-submitted websites and generate detailed reports on technical SEO issues, broken links, page speed, meta tags, and accessibility, feeding data into the Sitebeam SaaS platform. The bot was first documented in public forums around 2011 and has been consistently updated, with official documentation available at sitebeam.net/crawler.htm.

🌐 Technical Behavior

The Sitebeam crawler uses a modified webkit-based engine (Safari-like rendering) to simulate real user browsing, executing JavaScript and CSS to capture fully rendered page content. It sends a configurable number of concurrent requests, typically between 5 and 20 simultaneous connections, with a default crawl delay of 1–2 seconds between requests, though this can be adjusted by site owners via the Sitebeam control panel. IP ranges are not publicly documented but are known to originate from UK-based data centers (e.g., OVH, DigitalOcean) and may rotate. The bot uses HTTP/1.1 and supports both IPv4 and IPv6, sending a persistent Accept-Language: en-US,en;q=0.9 header. It follows all noindex and nofollow meta directives and does not index or store images beyond their URLs for alt-text analysis.

📋 robots.txt Compliance

Sitebeam fully honors robots.txt Disallow directives as confirmed in its official documentation and community reports. It also respects Crawl-Delay directives and can be blocked by adding a Disallow: / line for the User-Agent Sitebeam. The crawler checks robots.txt before each session and caches the file for up to 24 hours. There are no known incidents of its ignoring robots.txt rules.

🔍 Detection Indicators

The primary User-Agent string is Mozilla/5.0 (compatible; Sitebeam/2.0; +http://www.sitebeam.net/crawler.htm). Secondary variants include Sitebeam/1.0 (compatible; +http://www.sitebeam.net/crawler.htm). The bot consistently includes the Referer header set to the customer’s dashboard URL and does not send typical browser headers like Sec-Ch-Ua. Behavioral fingerprints include high request rates to CSS and JS assets, lack of mouse movement or scroll events, and sequential GET requests to 50–100 pages in under two minutes.

📊 Data Usage

Collected data is used solely for the Sitebeam SEO audit platform: generating reports on broken links (404 errors), duplicate content, page speed (using Lighthouse scores), missing meta descriptions, header structure, and schema markup. No data is sold to third parties or used for AI training. Customers access aggregated reports via a web dashboard; raw crawl logs are retained for 30 days and then deleted per GDPR-compliant data retention policies.

⚙️ Rate Limiting Policy

Sitebeam is rate-limited because its concurrent connection model can overwhelm smaller servers if left unchecked—threshold-based blocking (e.g., 50 requests per minute from a single IP) is recommended to prevent resource exhaustion while still allowing legitimate SEO audits. The official documentation advises contacting Sitebeam support for custom crawl speed adjustments if per-site rate limits are needed.

Free Traffic Analysis

What's Actually Crawling Your Website?

Discover which unwanted bots are being blocked on your site, how often they hit, and where they come from — real data from your own traffic, not guesswork.

🔍 Scan My Site Free

Powered by JA4 fingerprinting, honeypot traps & behavioral analysis

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.