ai2bot-deepresearcheval
AI2Bot-DeepResearchEval is a dedicated web crawler developed and operated by the Allen Institute for AI (AI2), a non-profit research institute based in Seattle. First publicly documented in early 2024, this bot is specifically designed to collect publicly accessible web content for the purpose of building evaluation datasets used in deep research benchmarking and to support training of AI2’s open‑source language models, such as OLMo and Tulu. Its creation was motivated by the need for high‑quality, verifiable web data to assess and improve the performance of large language models on complex reasoning and multi‑hop retrieval tasks, as described in AI2’s official publication (Allen AI, 2024).
The crawler employs a polite crawling policy with a default request rate of 1 request per 2 seconds per domain, though it may burst slightly during initial discovery phases. It uses HTTP/1.1 and HTTP/2 protocols, and its IP addresses are drawn from the ASN AS398963 (AI2 network) and a documented range of 204.14.232.0/21 (verified via whois and AI2’s published netblocks). AI2Bot-DeepResearchEval crawls primarily HTML pages, PDF documents, and JSON‑LD structured data, ignoring media files like images and videos. It does not follow JavaScript‑heavy redirects and limits its crawl depth to 3 hops from the seed list. The bot sets a non‑standard header X-AI2-Crawler: DeepResearchEval in addition to the User‑Agent string.
AI2 officially states that AI2Bot-DeepResearchEval fully obeys robots.txt directives, including Disallow and Crawl-Delay rules, as documented on their public crawler policy page at allenai.org/crawler. The bot also respects the User‑Agent: AI2Bot-DeepResearchEval token and will not crawl any path that is disallowed. Website operators can block the bot entirely by adding a Disallow: / line under the respective User‑Agent.
The primary identification is the User‑Agent string AI2Bot-DeepResearchEval/1.0 (https://allenai.org/crawler). Secondary fingerprints include the presence of the X-AI2-Crawler header and a request coming from an IP within the 204.14.232.0/21 range. The bot does not use a residential proxy network; all requests originate from AI2’s own data center IPs. Reverse DNS lookups on these IPs typically resolve to *.crawl.allenai.org.
Collected data is used exclusively for research evaluation and model training at AI2. The extracted text and metadata feed into curated datasets such as DeepResearchEvalBench, a benchmark released in June 2024 that tests a model’s ability to answer multi‑step research questions. AI2 also uses the data to fine‑tune retrieval‑augmented generation pipelines and to measure the factual accuracy of open‑source models. No collected content is sold or shared with third parties outside AI2’s research team.
Although entirely legitimate, AI2Bot-DeepResearchEval is rate‑limited by many web applications because its high‑volume, systematic crawl can consume significant server resources. Website operators typically impose a threshold‑based block (e.g., more than 50 requests per minute) to prevent performance degradation while still allowing the bot to complete its research mission.
Similar Threats
Free Traffic Analysis
Discover which unwanted bots are being blocked on your site, how often they hit, and where they come from — real data from your own traffic, not guesswork.
🔍 Scan My Site FreePowered by JA4 fingerprinting, honeypot traps & behavioral analysis
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.
Stay up to date with the latest from Boteraser.
We use cookies to improve your experience on our site. By using our site, you consent to cookies.
Manage your cookie preferences below:
Essential cookies enable basic functions and are necessary for the proper function of the website.
CloudFlare provides web performance and security solutions, enhancing site speed and protecting against threats.
Service URL: developers.cloudflare.com (opens in a new window)
These cookies are needed for adding comments on this website.
These cookies are used for managing login functionality on this website.
Statistics cookies collect information anonymously. This information helps us understand how visitors use our website.
Google Analytics is a powerful tool that tracks and analyzes website traffic for informed marketing decisions.
Service URL: policies.google.com (opens in a new window)
You can find more information in our Cookie Policy and Privacy Policy.