AI2Bot

Bot User-Agent: ai2bot

🤖 Overview

AI2Bot is a web crawler operated by the Allen Institute for AI (AI2), a nonprofit research institute founded by Microsoft co-founder Paul Allen. Officially documented in the institute’s crawler information page and its GitHub repositories, the bot’s primary purpose is to collect publicly available web content—particularly scientific articles, research papers, and academic datasets—to train and improve AI2’s open‑source language models, such as OLMo (Open Language Model), and to power the Semantic Scholar academic search engine. AI2 explicitly states the crawler is used solely for non‑commercial, research‑focused AI training and indexing.

🌐 Technical Behavior

AI2Bot employs a crawl strategy that prioritizes high‑quality, openly licensed content, with a strong emphasis on scholarly and educational domains (e.g., .edu, .org). According to the official documentation, the bot issues requests at a default rate of one request per second per host, but this can increase when crawling large repositories. It uses HTTP/1.1 and HTTPS, and its IP ranges are documented as belonging to the Amazon Web Services (AWS) cloud infrastructure, specifically the us‑west‑2 region. The bot always identifies itself via the User-Agent header and includes a crawl‑delay directive that can be customized in robots.txt. AI2Bot also respects noindex meta tags and robots.txt crawl‑delay values.

📋 robots.txt Compliance

AI2Bot fully honors robots.txt Disallow directives, as confirmed by AI2’s published crawl policy at https://allenai.org/robots.txt (which itself blocks certain paths). The institute provides a dedicated User-agent: AI2Bot section in their own robots.txt and explicitly instructs webmasters to use standard Disallow or Allow rules. Documentation notes that the bot will cease crawling any URL listed under Disallow within 24 hours of a robots.txt update, and it supports the Allow directive for granular control.

🔍 Detection Indicators

The primary User‑Agent string is Mozilla/5.0 (compatible; AI2Bot; +https://allenai.org/crawler.html). Additionally, the bot may use a secondary string AI2Bot/1.0 when making HEAD requests. Behavioral fingerprints include a consistent crawl delay of 1 second by default, and a tendency to request robots.txt before any other resource on a host. The bot also sets a From header with the email [email protected] for contact purposes. No known CVE entries exist for AI2Bot, as it is a benign research crawler.

📊 Data Usage

Collected data is used exclusively for non‑commercial AI research, including training of the OLMo series of large language models, refining the Semantic Scholar academic search index, and creating open‑source scientific corpora such as AI2’s Dolma dataset (a 3‑trillion‑token open corpus). All data is hosted on AI2’s own compute infrastructure and is not sold or shared with third parties for profit, as per their Crawling Policy.

⚙️ Rate Limiting Policy

AI2Bot is rate‑limited because its crawl frequency, while polite (1 req/s), can still generate significant load on smaller websites that host thousands of scholarly articles. Threshold‑based blocking (e.g., >10 req/s from its IP range) is justified to protect site stability, and AI2 encourages webmasters to set a Crawl-Delay directive in robots.txt if the default rate is too aggressive for their infrastructure.

Free Traffic Analysis

What's Actually Crawling Your Website?

Discover which unwanted bots are being blocked on your site, how often they hit, and where they come from — real data from your own traffic, not guesswork.

🔍 Scan My Site Free

Powered by JA4 fingerprinting, honeypot traps & behavioral analysis

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.