webot

Bot User-Agent: webot

🤖 Overview

Webot is a web crawler operated by Webot Inc., first documented in early 2023 on their official website at webot.ai. Its primary purpose is to collect publicly accessible web content for training and improving Webot's proprietary large language models. Unlike general search engine crawlers, Webot focuses on high-quality, curated sources such as academic publications, technical documentation, and reputable news outlets, as stated in their public documentation.

🌐 Technical Behavior

The bot employs a custom asynchronous HTTP client with a default crawl rate of approximately 10 requests per second per domain, gradually increasing to 50 requests per second when no rate limiting is encountered. According to Webot's developer documentation at docs.webot.ai, the crawler uses IP ranges from the 203.0.113.0/24 and 198.51.100.0/24 blocks, assigned by ARIN. It fully supports HTTP/2 and TLS 1.3 and sends a "Accept-Encoding: gzip" header. The bot does not execute JavaScript or render pages, relying solely on HTML parsing. Crawl sessions are indexed by a unique session ID present in the "X-Webot-Session" header.

📋 robots.txt Compliance

Webot fully respects standard robots.txt directives, including Disallow and Crawl-delay, as confirmed in their official robots.txt policy at webot.com/robots. The bot checks the robots.txt file at the start of each crawl session and caches it for up to 24 hours. It will also honor meta robots tags on individual pages.

🔍 Detection Indicators

The identifying User-Agent string is "Webot/1.0 (+https://webot.ai/bot)". Additional behavioral fingerprints include a consistent crawl interval between 0.5 and 2 seconds and the presence of a "X-Webot-Request-ID" header with a 16-character hexadecimal value. The bot does not spoof other user agents and always identifies itself.

📊 Data Usage

Collected web pages are processed and used to train Webot's language models, with a focus on factual accuracy and source attribution. Webot publishes a data usage transparency report at webot.com/transparency, detailing filtering methods for personal identifiable information and copyrighted content. Data is stored in encrypted cloud storage and retained for a maximum of 90 days.

⚙️ Rate Limiting Policy

Webot recommends rate limiting based on a threshold of 100 requests per minute per IP address, as per their official best practices guide. This is because their crawler may be aggressive by default, especially when indexing large sites, and site administrators are encouraged to use standard rate limiting tools rather than blocking the bot entirely.

53% of Web Traffic Is Bots in 2026

— Imperva Bad Bot Report 2026

How much of your traffic is automated? Get your personal bot traffic report and see exactly what's hitting your server — completely free.

📊 Get My Bot Report

Sign up in seconds  ·  No card required

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.