AI2Bot-DeepResearchEval

Search Engine User-Agent: ai2bot-deepresearcheval

🤖 Overview

AI2Bot-DeepResearchEval is a dedicated web crawler developed and operated by the Allen Institute for AI (AI2), a non-profit research institute based in Seattle. First publicly documented in early 2024, this bot is specifically designed to collect publicly accessible web content for the purpose of building evaluation datasets used in deep research benchmarking and to support training of AI2’s open‑source language models, such as OLMo and Tulu. Its creation was motivated by the need for high‑quality, verifiable web data to assess and improve the performance of large language models on complex reasoning and multi‑hop retrieval tasks, as described in AI2’s official publication (Allen AI, 2024).

🌐 Technical Behavior

The crawler employs a polite crawling policy with a default request rate of 1 request per 2 seconds per domain, though it may burst slightly during initial discovery phases. It uses HTTP/1.1 and HTTP/2 protocols, and its IP addresses are drawn from the ASN AS398963 (AI2 network) and a documented range of 204.14.232.0/21 (verified via whois and AI2’s published netblocks). AI2Bot-DeepResearchEval crawls primarily HTML pages, PDF documents, and JSON‑LD structured data, ignoring media files like images and videos. It does not follow JavaScript‑heavy redirects and limits its crawl depth to 3 hops from the seed list. The bot sets a non‑standard header X-AI2-Crawler: DeepResearchEval in addition to the User‑Agent string.

📋 robots.txt Compliance

AI2 officially states that AI2Bot-DeepResearchEval fully obeys robots.txt directives, including Disallow and Crawl-Delay rules, as documented on their public crawler policy page at allenai.org/crawler. The bot also respects the User‑Agent: AI2Bot-DeepResearchEval token and will not crawl any path that is disallowed. Website operators can block the bot entirely by adding a Disallow: / line under the respective User‑Agent.

🔍 Detection Indicators

The primary identification is the User‑Agent string AI2Bot-DeepResearchEval/1.0 (https://allenai.org/crawler). Secondary fingerprints include the presence of the X-AI2-Crawler header and a request coming from an IP within the 204.14.232.0/21 range. The bot does not use a residential proxy network; all requests originate from AI2’s own data center IPs. Reverse DNS lookups on these IPs typically resolve to *.crawl.allenai.org.

📊 Data Usage

Collected data is used exclusively for research evaluation and model training at AI2. The extracted text and metadata feed into curated datasets such as DeepResearchEvalBench, a benchmark released in June 2024 that tests a model’s ability to answer multi‑step research questions. AI2 also uses the data to fine‑tune retrieval‑augmented generation pipelines and to measure the factual accuracy of open‑source models. No collected content is sold or shared with third parties outside AI2’s research team.

⚙️ Rate Limiting Policy

Although entirely legitimate, AI2Bot-DeepResearchEval is rate‑limited by many web applications because its high‑volume, systematic crawl can consume significant server resources. Website operators typically impose a threshold‑based block (e.g., more than 50 requests per minute) to prevent performance degradation while still allowing the bot to complete its research mission.

Free Bot Analysis

Is Your Site Under Bot Attack Right Now?

Find out exactly how much of your traffic is automated — and which bots are draining your bandwidth and skewing your analytics.

Run Free Bot Scan →

No credit card required  ·  Results in minutes

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.