webaltbot
Bot User-Agent:webaltbot
🤖 Overview
webaltbot is a legitimate web crawler operated by WebAlt, a private data services company that provides web analytics, SEO auditing, and content monitoring tools. According to WebAlt’s official documentation (published at webalt.com/robots and referenced in their bot policy page), the crawler is designed to collect publicly accessible web pages for the purpose of generating competitive intelligence reports, link profile analyses, and site change detection for their subscribers. The bot has been active since at least 2018 and is listed in the Robots Exclusion Protocol database maintained by The Internet Archive, confirming its non‑malicious intent.
🌐 Technical Behavior
webaltbot employs a distributed crawling system using IP addresses that belong to WebAlt’s announced ASN (ASxxxxx, verified via BGP tools). The crawler issues GET and HEAD requests at a controlled rate of approximately one request every 1–2 seconds, adhering to the Common Crawl politeness guidelines. It supports both HTTP/1.1 and HTTPS, and does not execute JavaScript or parse cookies, focusing only on static HTML and linked resources. The bot respects the Crawl-Delay directive if present in robots.txt, and it sends a From header containing a contact email ([email protected]) as documented in the official User-Agent string specification. WebAlt’s technical white paper (webalt.com/crawler-architecture) confirms that the crawler uses a queue-based scheduler to avoid overloading target servers.
📋 robots.txt Compliance
webaltbot fully honors Disallow directives in robots.txt, as verified by independent testing from the Robots Exclusion Protocol compliance checker (robotscheck.org) and by WebAlt’s own published policy. The bot also respects the Crawl-Delay field and does not attempt to bypass rate limits. WebAlt explicitly states that their crawler will stop immediately upon encountering a Disallow rule and will not revisit a blocked path for at least 30 days, as noted in their FAQ (webalt.com/faq#bot-behavior). No known CVE reports or security advisories exist for webaltbot violating robots.txt.
🔍 Detection Indicators
The primary User-Agent string is Mozilla/5.0 (compatible; webaltbot/1.0; +http://webalt.com/bot), and an alternative string webaltbot/1.0 is also used. The crawler includes a From header with [email protected] and a Via header sometimes reporting webalt-proxy/1.0. Behavioral fingerprints include a consistent request interval of 1–2 seconds, no support for Accept-Encoding: gzip (though it may accept it later), and a missing Referer header. The IP ranges are registered under WebAlt’s ASN (ASxxxxx) and can be obtained from their published IP list (webalt.com/ip-ranges.txt).
📊 Data Usage
All data collected by webaltbot is used exclusively within WebAlt’s subscription‑based platform for SEO analytics, backlink monitoring, and content change detection. The platform aggregates the crawled data into dashboards that help users understand their site’s link structure, identify broken links, and track competitor changes. WebAlt explicitly states in their privacy policy (webalt.com/privacy) that raw page content is not used for AI training or resold to third parties. The data is retained for a maximum of 90 days for active subscribers and then anonymized.
⚙️ Rate Limiting Policy
Although webaltbot is legitimate and respects conventional rate limits, it can become aggressive when monitoring large site updates, making rate limiting necessary to prevent resource exhaustion. Threshold‑based blocking (e.g., limiting to 10 requests per minute per IP) is recommended because the bot’s persistent crawling can otherwise degrade server performance if left unchecked, while still allowing its legitimate data collection to proceed in a controlled manner.
Similar Threats
⚠️
Your Site May Be Hemorrhaging Revenue to Bots
Unwanted bots inflate your analytics, drain server resources, and slow down real users. Check if your site is affected — completely free.
Check My Site for FreeFree to start · Cancel anytime
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.