linkcheck by siteimprove com
Bot User-Agent:linkcheck-by-siteimprove-com
🤖 Overview
linkcheck by siteimprove com is a legitimate web crawler operated by Siteimprove, a Danish company founded in 2003 that provides digital governance and web quality assurance tools. The bot’s primary purpose is to systematically scan websites for broken links, spelling errors, accessibility violations (WCAG standards), SEO issues, and content quality problems, feeding data into the Siteimprove platform for customer dashboards and reports. According to Siteimprove’s official documentation (siteimprove.com/en/help/link-checking/), the crawler is part of their Content Quality and Accessibility modules, used by over 7,000 organizations worldwide including universities, government agencies, and enterprises.
🌐 Technical Behavior
The crawler uses a multithreaded, asynchronous architecture to scan pages at a high request rate, typically sending one request per second per thread but with configurable concurrency (default 10–20 threads). It follows HTTP/1.1 and HTTPS protocols, respecting Cache-Control and ETag headers to reduce server load. IP ranges are dynamic but frequently belong to AWS (Amazon Web Services) or Siteimprove’s own infrastructure; example netblocks include 52.84.0.0/15 and 54.239.0.0/16 based on Siteimprove’s published IP lists (siteimprove.com/en/support/faq/). The bot performs full-site crawls at scheduled intervals (daily, weekly, or monthly depending on customer plan) and can also be triggered manually. It identifies itself via the Accept-Language: en-US header and requests User-Agent: Mozilla/5.0 (compatible; linkcheck by siteimprove.com; +https://siteimprove.com/en/support/link-checking/). The crawler does not execute JavaScript or render pages; it only parses static HTML and HTTP responses.
📋 robots.txt Compliance
Siteimprove’s official documentation explicitly states that linkcheck by siteimprove com fully honors robots.txt Disallow directives and the Crawl-Delay meta tag. The crawler reads the robots.txt file at the start of each crawl and will skip URLs blocked by rules. However, Siteimprove recommends that customers who wish to exclude specific paths use the Disallow: /path directive for the User-agent: linkcheck by siteimprove.com line. Evidence from Siteimprove’s support articles (siteimprove.com/en/support/how-to-control-link-checker-crawls/) confirms that the bot checks robots.txt before each crawl session and during incremental scans.
🔍 Detection Indicators
The primary User-Agent string is Mozilla/5.0 (compatible; linkcheck by siteimprove.com; +https://siteimprove.com/en/support/link-checking/). A secondary string Siteimprove-Link-Checker/1.0 may appear in some legacy deployments. Behavioral fingerprint: the bot issues a high volume of requests in short bursts (often 10–20 requests per second) but respects server response times and will slow down if the server returns 503 or 429 status codes. It does not submit forms, click links, or interact with cookies beyond session handling. Identifying headers include X-Siteimprove-Crawler: 1 in some newer versions, as documented in Siteimprove’s public knowledge base.
📊 Data Usage
Collected data is used exclusively for Siteimprove’s web governance platform, providing customers with actionable reports on broken links (HTTP 404/410), misspellings, accessibility violations (e.g., missing alt text, contrast issues), and SEO metadata quality. The data is stored in Siteimprove’s cloud infrastructure and is not used for AI training, advertising, or resale. Siteimprove’s privacy policy (siteimprove.com/en/privacy/) confirms that crawled content is retained only as long as necessary for the customer’s subscription and is not shared with third parties.
⚙️ Rate Limiting Policy
While not malicious, linkcheck by siteimprove com can be aggressive during initial full-site crawls, issuing up to 20 requests per second per thread. It is rate-limited (e.g., blocked after 50 requests per second) to prevent server overload and to protect web application performance. Siteimprove advises customers to set Crawl-Delay: 10 in robots.txt if they need to throttle the bot further. This policy is based on industry-standard threshold-based blocking (e.g., Nginx `limit_req_zone`) to balance thorough scanning with server resource availability.
Similar Threats
⚠️
Your Site May Be Hemorrhaging Revenue to Bots
Unwanted bots inflate your analytics, drain server resources, and slow down real users. Check if your site is affected — completely free.
Check My Site for FreeFree to start · Cancel anytime
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.