ExaBot
Bot User-Agent:exabot
🤖 Overview
ExaBot is a web crawler operated by Exa (exa.ai), a company that provides a semantic search engine for AI and enterprise applications, launched in 2021. It collects publicly available web pages to build Exa's neural retrieval index, which powers their search API and supports AI training datasets focused on contextual meaning rather than keywords.
🌐 Technical Behavior
ExaBot uses IP address ranges officially listed on Exa's documentation page at https://docs.exa.ai/crawler/ip-ranges and in their GitHub repository at https://github.com/exa-labs/crawler-info. It operates over HTTP/1.1 and HTTP/2, with a default crawl interval of 2–5 seconds per domain, but can accelerate to 10 requests per second for responsive sites. The crawler respects Cache-Control and If-Modified-Since headers, follows robots.txt directives including Crawl-Delay, and parses sitemaps for prioritization. ExaBot does not execute JavaScript or CSS, focusing solely on raw HTML and visible text content.
📋 robots.txt Compliance
Exa's official robots policy at https://exa.ai/robots confirms that ExaBot fully adheres to the Robots Exclusion Protocol. It honors Disallow and Allow directives, as well as the X-Robots-Tag HTTP header for per-URL control. No evidence of violations has been reported in security advisories or community discussions.
🔍 Detection Indicators
The primary User-Agent string is ExaBot/1.0 (https://exa.ai/crawler); a variant ExaBot/2.0 exists for updated versions. The crawler sends standard HTTP headers (Accept, Accept-Encoding) and does not impersonate browsers. Behaviorally, ExaBot tends to crawl pages with high textual density—scientific articles, long-form writing—preferring off-peak hours (UTC 00:00–06:00).
📊 Data Usage
Data collected is used exclusively for Exa's semantic vector search index, where content is transformed into embeddings for context-aware querying. Exa's privacy policy states that crawled data is not sold to third parties nor used to train external large language models. Instead, it fuels Exa's own AI search API, which serves researchers, data scientists, and enterprise teams.
⚙️ Rate Limiting Policy
Although ExaBot is a legitimate, well-behaved crawler, its aggressive pursuit of fresh content—especially on news and academic domains—can cause load. A rate-limiting threshold of 100 requests per minute is strongly recommended to protect server resources while enabling periodic re-crawls.
Similar Threats
Free Traffic Analysis
What's Actually Crawling Your Website?
Discover which unwanted bots are being blocked on your site, how often they hit, and where they come from — real data from your own traffic, not guesswork.
🔍 Scan My Site FreePowered by JA4 fingerprinting, honeypot traps & behavioral analysis
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.