xaldon
Bot User-Agent:xaldon
🤖 Overview
Xaldon is a web crawler operated by Xaldon Global, a company providing web monitoring and digital asset intelligence services, primarily used for competitive analysis, brand protection, and SEO auditing. According to Xaldon’s official website (xaldon.com), the bot collects publicly accessible web content to feed the company’s Digital Asset Intelligence Platform, which helps clients track changes in competitor web properties, detect unauthorized use of brand assets, and monitor domain ownership shifts.
🌐 Technical Behavior
Xaldon crawls using HTTP/1.1 and HTTP/2 protocols with a default crawl rate of approximately 1 request per 2 seconds from a single IP, scaling to up to 5 requests per second when multiple IPs are used concurrently. According to Xaldon’s published IP ranges (as of January 2025, listed in their official documentation at docs.xaldon.com/crawler), the bot originates from IP blocks within the 2a06:6300::/29 (IPv6) and 185.93.0.0/22 (IPv4) ranges. It respects ETag and Last-Modified headers to reduce redundant downloads, but does not cache content across multiple domains. The crawler uses a depth-first traversal strategy, starting from sitemaps specified in robots.txt or discovered via meta tags, and can follow up to 3 levels of internal links per session before moving to a new seed URL.
📋 robots.txt Compliance
Xaldon is documented to honor Disallow directives in robots.txt, as stated in its official guidelines (xaldon.com/robots). However, there are reported instances (e.g., a 2023 GitHub issue on xaldon/crawler-support) where it ignored Crawl-delay directives, leading to community requests for stricter rate limiting. Xaldon Global responded by updating their crawler to respect crawl-delay in v2.1 released in March 2024. It also supports Allow overrides under Disallow patterns.
🔍 Detection Indicators
The primary User-Agent string is Mozilla/5.0 (compatible; Xaldon/2.1; +https://xaldon.com/bot). Older versions used a similar string without the version number. Xaldon also sends a custom HTTP header X-Xaldon-Crawl-ID containing a UUID for each crawl session, and includes a Referer header pointing to the previous page crawled. Behavioral fingerprinting shows it always requests robots.txt before any other resource on a new domain, and it does not fetch images or CSS unless specifically allowed via a Xaldon-Fetch-Assets header.
📊 Data Usage
Collected data, including full HTML pages, meta tags, link structures, and response headers, is ingested into Xaldon’s platform for web change detection, backlink analysis, and brand compliance monitoring. A 2024 whitepaper from Xaldon (available at xaldon.com/research/whitepaper-2024) states that content is not used for AI training or general search indexing, but strictly for client-specific analytics. Data is retained for a maximum of 90 days, after which it is aggregated or deleted.
⚙️ Rate Limiting Policy
While Xaldon is legitimate and adheres to robots.txt, its moderate crawl rate and occasional non-compliance with historical Crawl-delay directives justify rate limiting as a precautionary measure to prevent server overload. A threshold-based block (e.g., >10 requests per 10 seconds from a single IP) is recommended to protect application responsiveness without impeding legitimate monitoring.
Similar Threats
Free Bot Analysis
Is Your Site Under Bot Attack Right Now?
Find out exactly how much of your traffic is automated — and which bots are draining your bandwidth and skewing your analytics.
Run Free Bot Scan →No credit card required · Results in minutes
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.