SiteCheckerBotCrawler

Crawler User-Agent: sitecheckerbotcrawler

🤖 Overview

SiteCheckerBotCrawler is a legitimate web crawler operated by SiteChecker (sitechecker.pro), a website monitoring and SEO auditing platform headquartered in the United States. Its primary purpose is to systematically scan websites to collect data on uptime, page load speed, SEO health, broken links, accessibility issues, and security vulnerabilities, feeding results into the SiteChecker dashboard for website owners and administrators. The bot was first documented in 2018 and is widely recognized in the web monitoring community for its systematic crawling behavior.

🌐 Technical Behavior

The crawler initiates scans from a set of IP addresses primarily belonging to Amazon Web Services (AWS) and DigitalOcean, with ranges such as 54.236.1.0/24 and 159.89.0.0/16 frequently observed in logs. It employs HTTP/1.1 and HTTP/2 protocols, sending requests with a User-Agent string of Mozilla/5.0 (compatible; SiteCheckerBotCrawler/1.0; +https://sitechecker.pro/bot). By default, it crawls at a rate of approximately 5–10 requests per second per domain, though this can be adjusted by website operators via the SiteChecker dashboard. The bot respects robots.txt directives and will not crawl paths marked as Disallow; it also supports Crawl-Delay header values to throttle itself. It retrieves content via GET and HEAD methods, parsing HTML, CSS, JavaScript, and images to evaluate page performance. According to SiteChecker’s official documentation, the crawler does not execute JavaScript for dynamic content analysis but focuses on static resource retrieval.

📋 robots.txt Compliance

Based on SiteChecker’s published guidelines at https://sitechecker.pro/robots.txt, the SiteCheckerBotCrawler fully honors Disallow directives in standard robots.txt files. Network administrators have independently verified this by observing that the bot ceases crawling on paths such as /wp-admin or /admin when those are disallowed. No public CVE entries or security advisories have been filed against the bot for violating robots.txt exclusion.

🔍 Detection Indicators

The definitive User-Agent string is SiteCheckerBotCrawler/1.0 without additional environment spoofing. Behavioral fingerprints include sequential page crawling with consistent intervals, absence of JavaScript execution, and an IP origin from AWS or DigitalOcean. Administrators can also check for an X-SiteChecker-Request header (value true) that the bot attaches to all its requests, a feature documented in SiteChecker’s API reference at https://sitechecker.pro/docs/api.

📊 Data Usage

Data collected by the bot is used exclusively for the SiteChecker platform’s core features: website uptime monitoring, page speed scoring (informed by Google Lighthouse metrics), broken link detection, SEO issue reporting, and accessibility compliance checks. The aggregated results are presented to the website owner through the SiteChecker dashboard; no data is sold to third parties or used for AI training, as confirmed by SiteChecker’s privacy policy at https://sitechecker.pro/privacy.

⚙️ Rate Limiting Policy

SiteCheckerBotCrawler is rate-limited because its standard crawl frequency of 5–10 requests per second can strain shared hosting environments or trigger server‑side alarms if not constrained. The recommended threshold for blocking without harming legitimate monitoring is to set a rate limit of 50 requests per minute per IP, with a 15‑second pause after hitting the limit, as advised in SiteChecker’s support documentation at https://sitechecker.pro/support/rate-limiting.

53% of Web Traffic Is Bots in 2026

— Imperva Bad Bot Report 2026

How much of your traffic is automated? Get your personal bot traffic report and see exactly what's hitting your server — completely free.

📊 Get My Bot Report

Sign up in seconds  ·  No card required

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.