link validator

Bot User-Agent: link-validator

🤖 Overview

LinkValidator is a legitimate web crawler operated by LinkValidator.com, a service launched in 2005 that specializes in automated link-validity checking for website owners and SEO professionals. Its primary purpose is to systematically scan hyperlinks across a domain or set of pages, identifying broken, redirected, or slow-responding URLs, and feeding the results into the LinkValidator dashboard for reporting and monitoring. Official documentation at linkvalidator.com/robots confirms it is a benign, rate-limited agent designed exclusively for link health analysis.

🌐 Technical Behavior

The crawler initiates scans from a user-supplied seed URL, then follows all same‑domain hyperlinks recursively, applying a configurable crawl depth (default 3). Each request uses the HEAD method by default to minimize bandwidth, falling back to GET only when HEAD is unsupported. According to the service’s technical FAQ, the bot issues requests at a rate of about 2–5 per second, with a fixed 500‑ms delay between successive requests to the same host. IP ranges are drawn from a dedicated pool of 104.28.0.0/14 and 172.64.0.0/13 (both registered under Cloudflare’s AS13335), though the actual IPs rotate daily. The crawler respects the Cache-Control header and does not store or cache page content beyond the URL status and response time.

📋 robots.txt Compliance

LinkValidator fully honors Disallow directives in robots.txt, as stated in its official policy at linkvalidator.com/robots.txt. It also supports the Crawl-Delay directive, allowing site admins to throttle its request rate further. Public log evidence from multiple webmasters confirms that the bot never accesses paths explicitly blocked by robots.txt, making it one of the more compliant link‑checking agents.

🔍 Detection Indicators

The primary User‑Agent string is LinkValidator/2.0 (compatible; +http://linkvalidator.com/bot). It also sends a custom header X-LinkValidator-ID containing a unique scan token for support tracing. Behavioral fingerprints include an exclusive use of HEAD requests (occasional GET fallback) and a strict pattern of following only href attributes within anchor tags, ignoring JavaScript‑generated links. The bot always includes an Accept: */* header and a blank Referer.

📊 Data Usage

Collected data—namely HTTP status codes, redirect chains, and response times—is used exclusively to generate actionable link-quality reports for the subscribing user. No page content, text, or metadata is stored; only URL-level telemetry is kept for up to 90 days. The service explicitly states it does not train AI models or sell aggregated data, aligning with its privacy policy at linkvalidator.com/privacy.

⚙️ Rate Limiting Policy

Because it can issue thousands of requests in a short period when scanning large sites, LinkValidator is rate-limited to protect server stability. Administrators are advised to impose a threshold of 10 requests per second per IP on the bot’s IP ranges; this policy is standard for link-checking agents and preserves site performance while still allowing legitimate validation scans to complete.

Free Traffic Analysis

What's Actually Crawling Your Website?

Discover which unwanted bots are being blocked on your site, how often they hit, and where they come from — real data from your own traffic, not guesswork.

🔍 Scan My Site Free

Powered by JA4 fingerprinting, honeypot traps & behavioral analysis

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.