exabot
ExaBot is a web crawler operated by Exa (exa.ai), a company that provides a semantic search engine for AI and enterprise applications, launched in 2021. It collects publicly available web pages to build Exa's neural retrieval index, which powers their search API and supports AI training datasets focused on contextual meaning rather than keywords.
ExaBot uses IP address ranges officially listed on Exa's documentation page at https://docs.exa.ai/crawler/ip-ranges and in their GitHub repository at https://github.com/exa-labs/crawler-info. It operates over HTTP/1.1 and HTTP/2, with a default crawl interval of 2–5 seconds per domain, but can accelerate to 10 requests per second for responsive sites. The crawler respects Cache-Control and If-Modified-Since headers, follows robots.txt directives including Crawl-Delay, and parses sitemaps for prioritization. ExaBot does not execute JavaScript or CSS, focusing solely on raw HTML and visible text content.
Exa's official robots policy at https://exa.ai/robots confirms that ExaBot fully adheres to the Robots Exclusion Protocol. It honors Disallow and Allow directives, as well as the X-Robots-Tag HTTP header for per-URL control. No evidence of violations has been reported in security advisories or community discussions.
The primary User-Agent string is ExaBot/1.0 (https://exa.ai/crawler); a variant ExaBot/2.0 exists for updated versions. The crawler sends standard HTTP headers (Accept, Accept-Encoding) and does not impersonate browsers. Behaviorally, ExaBot tends to crawl pages with high textual density—scientific articles, long-form writing—preferring off-peak hours (UTC 00:00–06:00).
Data collected is used exclusively for Exa's semantic vector search index, where content is transformed into embeddings for context-aware querying. Exa's privacy policy states that crawled data is not sold to third parties nor used to train external large language models. Instead, it fuels Exa's own AI search API, which serves researchers, data scientists, and enterprise teams.
Although ExaBot is a legitimate, well-behaved crawler, its aggressive pursuit of fresh content—especially on news and academic domains—can cause load. A rate-limiting threshold of 100 requests per minute is strongly recommended to protect server resources while enabling periodic re-crawls.
Similar Threats
🛡️
Boteraser automatically detects and blocks unwanted bots — protecting your site from scrapers, DDoS bursts, and credential stuffing attacks without slowing down real visitors.
✅ Start Free ProtectionSetup takes under a minute · Free trial available
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.
Stay up to date with the latest from Boteraser.
We use cookies to improve your experience on our site. By using our site, you consent to cookies.
Manage your cookie preferences below:
Essential cookies enable basic functions and are necessary for the proper function of the website.
CloudFlare provides web performance and security solutions, enhancing site speed and protecting against threats.
Service URL: developers.cloudflare.com (opens in a new window)
These cookies are needed for adding comments on this website.
These cookies are used for managing login functionality on this website.
Statistics cookies collect information anonymously. This information helps us understand how visitors use our website.
Google Analytics is a powerful tool that tracks and analyzes website traffic for informed marketing decisions.
Service URL: policies.google.com (opens in a new window)
You can find more information in our Cookie Policy and Privacy Policy.