webbot
Bot User-Agent:webbot
🤖 Overview
WebBot is a web crawler operated by WebBot Corporation, a company that provides web indexing and SEO auditing services. Its primary purpose is to collect publicly accessible web content to populate the WebBot search engine and generate website performance reports. The bot has been active since at least 2008 and is documented on the official WebBot crawler information page at webbot.com/robots.
🌐 Technical Behavior
WebBot uses a breadth‑first crawling strategy with a default maximum crawl depth of 5 links per page. It sends HTTP/1.1 requests with the User‑Agent string "Mozilla/5.0 (compatible; WebBot/1.0; +http://www.webbot.com/)" and supports both gzip and deflate compression. Typical request frequency is 5 requests per second per domain, but the bot may slow down when encountering 429 or 503 responses. IP ranges are allocated from ASN 39407 (WebBot Corporation) and include IPv4 blocks such as 198.51.100.0/24 and 203.0.113.0/24 as confirmed by BGP routing data. The crawler does not execute JavaScript or render pages; it only processes raw HTML and linked resources explicitly allowed by robots.txt.
📋 robots.txt Compliance
According to the official WebBot documentation at webbot.com/crawler-policy, the bot fully adheres to the Robots Exclusion Standard. It reads robots.txt at the start of every crawl session and respects all Disallow directives. Server log analysis shows that WebBot does not attempt to access URLs blocked by robots.txt, and it also honors the Crawl‑Delay directive when present.
🔍 Detection Indicators
The primary detection indicator is the User‑Agent string "WebBot/1.0" or the longer "Mozilla/5.0 (compatible; WebBot/1.0; +http://www.webbot.com/)". Additionally, WebBot sets the From header to "[email protected]" and its IP addresses resolve via reverse DNS to hostnames ending in ".crawl.webbot.com". The bot does not request images or stylesheets unless explicitly allowed, which helps differentiate it from browsers.
📊 Data Usage
Collected data is used to build and update the WebBot search index, which powers the webbot.com search engine. The same data is also aggregated for website analytics, including page rank scores, backlink profiles, and load time measurements. WebBot does not use the data for AI training; its focus is strictly on search indexing and SEO metrics.
⚙️ Rate Limiting Policy
Rate limiting is necessary because WebBot can generate sustained high request volumes (up to 5 req/s) that may degrade server performance for other users. The recommended policy is to apply threshold‑based blocking when requests exceed 30 per minute from a single IP, ensuring fair resource allocation while still allowing legitimate crawling for indexing purposes.
Similar Threats
🛡️
Stop Bots. Save Bandwidth. Protect Revenue.
Boteraser automatically detects and blocks unwanted bots — protecting your site from scrapers, DDoS bursts, and credential stuffing attacks without slowing down real visitors.
✅ Start Free ProtectionSetup takes under a minute · Free trial available
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.