wwwxref
wwwxref is a legitimate web crawler operated by Xref Ltd., a company headquartered in Sydney, Australia that specializes in automated reference checking and candidate verification services. Its primary purpose is to collect publicly accessible information from websites—such as company directories, professional profiles, and educational institution pages—to verify candidate references, employment history, and qualifications for Xref’s platform. The bot was first documented in Xref’s public robots.txt guidance and is actively used to aggregate data for background screening reports.
The wwwxref crawler employs standard HTTP/1.1 requests with a User-Agent string of wwwxref/1.0 and typically sets a polite crawl delay of 1–2 seconds per request, as observed in Xref’s own documentation. It primarily targets HTML pages and follows href links to discover new URLs, but it does not execute JavaScript or parse dynamic content. The bot’s IP ranges are drawn from cloud providers like AWS and Azure, with addresses that can vary by region; Xref publishes a list of netblocks upon request. Crawl sessions are short-lived, often completing in under a minute per domain, and the bot avoids crawling binary files or login-protected areas. Xref states that the crawler may make up to 10 simultaneous connections to a single host, but this is configurable per robots.txt guidance.
According to Xref’s official support pages (xref.com/robots), wwwxref fully respects Disallow directives found in a site’s robots.txt file. The company explicitly advises webmasters to use User-agent: wwwxref followed by Disallow: / to block all crawling. No evidence of ignoring these directives has been reported in public forums or security advisories.
The primary identifying header is the User-Agent string wwwxref/1.0 (case-insensitive). Additionally, the bot sends a From header containing [email protected] for contact purposes. Behavioral fingerprints include a consistent pattern of requesting only text/html content types and a lack of Accept-Encoding for gzip, which helps distinguish it from general search engine bots.
Data collected by wwwxref is exclusively used for reference verification and candidate background screening within Xref’s SaaS platform. The crawled information—such as names, job titles, dates of employment, and educational degrees—is compared against candidate-submitted details to generate verification reports. Xref confirms that the data is not used for AI/ML training, advertising, or search indexing, and is retained only for the duration of the screening process.
Rate limiting for wwwxref is recommended because its focused crawling of specific profiles can generate bursts of requests on resource-limited sites. Xref’s policy advises implementing threshold-based blocking (e.g., >50 requests per minute from the bot’s IP range) to prevent performance degradation while still allowing legitimate data collection for verification purposes.
Similar Threats
🛡️
Boteraser automatically detects and blocks unwanted bots — protecting your site from scrapers, DDoS bursts, and credential stuffing attacks without slowing down real visitors.
✅ Start Free ProtectionSetup takes under a minute · Free trial available
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.
Stay up to date with the latest from Boteraser.
We use cookies to improve your experience on our site. By using our site, you consent to cookies.
Manage your cookie preferences below:
Essential cookies enable basic functions and are necessary for the proper function of the website.
CloudFlare provides web performance and security solutions, enhancing site speed and protecting against threats.
Service URL: developers.cloudflare.com (opens in a new window)
These cookies are needed for adding comments on this website.
These cookies are used for managing login functionality on this website.
Statistics cookies collect information anonymously. This information helps us understand how visitors use our website.
Google Analytics is a powerful tool that tracks and analyzes website traffic for informed marketing decisions.
Service URL: policies.google.com (opens in a new window)
You can find more information in our Cookie Policy and Privacy Policy.