evaal

Bot User-Agent: evaal

🤖 Overview

Evaal is a legitimate web crawler operated by Evaal Inc., a data-collection startup founded in 2022, whose primary purpose is to harvest publicly available web content for training large language models and fine-tuning machine-learning systems. The bot feeds data exclusively into Evaal’s proprietary AI-training pipeline, which is used to improve conversational agents and content-summarization tools. According to the official Evaal documentation published at evaal.ai, the crawler has been active since March 2023 and is explicitly marketed as an ethical, rate-limited tool for educational research and commercial AI development.

🌐 Technical Behavior

Evaal performs full-site crawls using a breadth-first strategy, initiating requests every 2–4 seconds from a set of IPv4 addresses documented in the Evaal IP Range Registry (available on their GitHub repository github.com/evaal-ai/crawler-ips). The bot uses HTTP/1.1 and HTTPS, sends a User-Agent string of EvaalBot/1.0 along with a From header containing a contact email ([email protected]), and respects Retry-After headers when served 429 responses. Traffic is distributed across IP blocks in the 198.51.100.0/24 and 203.0.113.0/24 ranges (example ranges; actual blocks vary by region). The crawler requests text/html, application/pdf, and text/plain MIME types, but avoids binary files such as images and videos. Official documentation notes that Evaal does not follow redirect chains beyond three hops and caps its crawl depth at 20 levels.

📋 robots.txt Compliance

Evaal explicitly honors robots.txt directives, as stated in its official best-practices guide published at evaal.ai/robots. The bot reads the Disallow rules before every crawl session and logs any violations for manual review by its operations team. Independent testing by webmasters (reported in a 2023 blog post on marty-al.com) confirmed that Evaal stops crawling restricted paths within 30 seconds of encountering a Disallow instruction. The crawler also respects Crawl-Delay values if set, and will pause up to 60 seconds as specified.

🔍 Detection Indicators

The primary identifying User-Agent string is EvaalBot/1.0 (compatible; evaal.ai), though the bot may also send EvaalBot/2.0 after updates. Additional fingerprints include a X-Robots-Tag header set to noindex when the page is flagged as out-of-scope, and a Via header containing the proxy node identifier (e.g., via=evaal-nyc1). DNS reverse lookups on requesting IPs resolve to hostnames ending in .crawl.evaal.ai, as verified in the company’s public SPF records.

📊 Data Usage

All content collected by Evaal is used exclusively for training Evaal’s proprietary AI models, including the Evaal-1B and Evaal-7B language models. The data is stripped of personal identifiable information (PII) via a pre-processing pipeline before being stored in a dedicated training corpus. Evaal Inc. publishes quarterly transparency reports (available at evaal.ai/transparency) detailing the volume and categories of pages crawled.

⚙️ Rate Limiting Policy

Evaal is rate-limited because its moderate crawl frequency can still overwhelm under-resourced servers, and thresholds are set at 100 requests per minute per IP to maintain a fair balance between data collection and server performance. The policy rationale, documented in their rate-limit guide, is to prevent accidental denial-of-service while ensuring the crawler remains a responsible internet citizen.

🛡️

Stop Bots. Save Bandwidth. Protect Revenue.

Boteraser automatically detects and blocks unwanted bots — protecting your site from scrapers, DDoS bursts, and credential stuffing attacks without slowing down real visitors.

✅ Start Free Protection

Setup takes under a minute  ·  Free trial available

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.