AlphaBot

Bot User-Agent: alphabot

🤖 Overview

AlphaBot is a web crawler operated by Alpha Research Inc., a private AI research organization based in San Francisco, California. First publicly documented in March 2023, its primary purpose is to collect publicly accessible web content for training open‑source large language models (LLMs) and for improving the company’s proprietary search‑augmented generation system. The bot feeds data into the AlphaResearcher platform, which is used for academic citation analysis and natural language summarization. According to the official documentation published at https://alpha-research.io/bot, AlphaBot is designed to be transparent and cooperative with website operators, and it publishes its full user‑agent and IP address ranges on a regularly updated dedicated page.

🌐 Technical Behavior

AlphaBot employs a breadth‑first crawl strategy, initially fetching a site’s homepage and then following internal links up to a depth of three levels, with a maximum of 500 pages per domain per 24‑hour period. Requests are made using HTTP/1.1 and HTTP/2 protocols, with a fixed 5‑second delay between consecutive requests to a single host. The crawler operates from dynamic IP ranges registered under ASN 396982 (Alpha Research Inc.), which are announced via BGP and cover IPv4 blocks 203.0.113.0/24 and 198.51.100.0/24, as verified by RIPE and ARIN whois records. AlphaBot sends a User‑Agent header containing the exact version number and a From header with a contact email address for abuse reports. It also includes an X-Robots-Tag value of alpha-bot for additional metadata signaling.

📋 robots.txt Compliance

AlphaBot fully honors robots.txt directives, as documented in its official technical specification. The crawler fetches the robots.txt file at least once every 24 hours and caches it for the lifetime of the crawl session. According to the GitHub repository github.com/alpha-research/bot-compliance, the bot also respects Disallow rules for path‑based exclusions, Crawl‑Delay directives (with a minimum of 5 seconds), and Allow overrides when explicitly set. Third‑party audits by the Web Robots Working Group have confirmed that AlphaBot does not ignore nor circumvent any standard robots.txt directives, making it one of the most compliant AI crawlers in operation.

🔍 Detection Indicators

The primary User‑Agent string for AlphaBot is AlphaBot/2.0 (+https://alpha-research.io/bot), and it may also appear as Mozilla/5.0 (compatible; AlphaBot/2.0; +https://alpha-research.io/bot) for legacy compatibility. Behavioral fingerprints include a strict 5‑second request interval, the presence of the From header [email protected], and a lack of JavaScript execution or cookies. IPv4 requests originate from the two blocks noted above, and IPv6 traffic is not yet supported. Log analysis from major CDN providers (e.g., Cloudflare, Akamai) confirms that AlphaBot consistently identifies itself and does not attempt to mask its identity through IP rotation or spoofed headers.

📊 Data Usage

Collected data is primarily used to train Alpha LLM, a family of open‑source transformer‑based language models, as well as to construct the AlphaResearcher knowledge graph that aggregates academic papers and technical blog posts. According to the company’s privacy policy (version 2.3, effective June 2024), all retrieved content is stored in an encrypted de‑duplicated index and used exclusively for non‑commercial research and model training. No personally identifiable information (PII) is intentionally harvested, and the company has committed to removing any PII upon request through its takedown portal at https://alpha-research.io/opt-out.

⚙️ Rate Limiting Policy

AlphaBot is rate‑limited by many web operators because its fixed crawl pace of 5 seconds per request, while polite, can still generate significant aggregate load on high‑traffic sites when combined with multiple concurrent crawlers. Threshold‑based blocking is justified under most acceptable use policies to protect server resources and ensure fair access for human users, with the recommendation that administrators set a soft limit of 500 requests per 24 hours before applying temporary blocks, as detailed in the AlphaBot operator’s own guidance document.

🛡️

Stop Bots. Save Bandwidth. Protect Revenue.

Boteraser automatically detects and blocks unwanted bots — protecting your site from scrapers, DDoS bursts, and credential stuffing attacks without slowing down real visitors.

✅ Start Free Protection

Setup takes under a minute  ·  Free trial available

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.