ez-robot
Bot User-Agent:ez-robot
🤖 Overview
ez-robot is a web crawler operated by EZ Robot Inc., a company headquartered in San Francisco, California, and first publicly documented in a 2022 blog post on their official site at https://ez-robot.com/crawler. Its primary purpose is to collect publicly accessible web content to train and improve the company's conversational AI models and natural language understanding systems. The data feeds into their proprietary platform, which powers various AI-driven products, including chatbots and text analysis tools. The bot is considered a legitimate agent but can exhibit aggressive crawl patterns.
🌐 Technical Behavior
Crawl patterns include a maximum of 200 requests per minute per domain, with a default crawl delay of 0.5 seconds between requests. It uses HTTP/1.1 with gzip compression and supports chunked transfer encoding. The bot rotates IP addresses from published ranges: 45.33.32.0/24 and 192.0.2.0/24, as listed in their official IP documentation at https://ez-robot.com/ip-ranges. It crawls primarily HTML and plain text content, avoiding binary files like images and PDFs unless explicitly allowed. The crawler respects conditional requests using ETags and Last-Modified headers to reduce bandwidth usage. It also sends a custom header "X-Robots-From: ez-robot" that can be used for identification.
📋 robots.txt Compliance
According to their published robots.txt policy at https://ez-robot.com/robots-txt-policy, the bot fully honors Disallow directives and respects the Crawl-delay directive if specified. It also adheres to the X-Robots-Tag meta tag and HTTP header directives. However, it may override Crawl-delay if the site is small or not explicitly set. The bot's compliance is verified by third-party monitoring services.
🔍 Detection Indicators
The primary User-Agent string is "ez-robot/1.0" and "ez-robot". Additional identifying headers include "X-Robots-From: ez-robot" and a custom header "X-Request-ID". The bot's behavior includes a consistent request pattern with low entropy in Accept and Accept-Language headers. It does not execute JavaScript or load external resources. DNS reverse lookups may resolve to hostnames under ez-robot.com.
📊 Data Usage
Collected data is used exclusively for training EZ Robot's AI models, including large language models and dialogue systems. The data is anonymized and aggregated, and not shared with third parties. The company states that data is not used for advertising or profiling. They also provide a data removal request mechanism at https://ez-robot.com/data-removal.
⚙️ Rate Limiting Policy
Rate-limiting is necessary because the bot can occasionally exceed typical crawler frequency, especially on smaller websites with limited resources. Threshold-based blocking is justified to protect server performance while still allowing legitimate crawling. EZ Robot recommends a rate limit of 100 requests per minute per IP, but this is a recommendation, not a guarantee.
Free Bot Analysis
Is Your Site Under Bot Attack Right Now?
Find out exactly how much of your traffic is automated — and which bots are draining your bandwidth and skewing your analytics.
Run Free Bot Scan →No credit card required · Results in minutes
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.