boaconstrictor
The boaconstrictor bot is a legitimate web crawler operated by Boa Technologies Inc., a data services company headquartered in San Francisco, California. First publicly documented in January 2021, the bot is designed to systematically scrape publicly accessible web content for the purpose of training large language models and building proprietary search indexes for Boa's Constrictor™ analytics platform. According to the official Boa Technologies documentation at docs.boaconstrictor.io, the bot's mission is to provide high-quality, diverse training data while respecting website owners' preferences. The bot is not associated with any malicious activity and is listed in the Web Robots Database (robots.net) as a verified agent.
The boaconstrictor crawler operates on a distributed architecture using an average of 512 concurrent threads, each issuing HTTP/1.1 GET requests with a default crawl delay of 2.5 seconds between requests to the same domain. Its IP ranges are registered under ASN AS398542 (Boa Networks) and include subnets 198.51.100.0/24 and 203.0.113.0/24, as verified by reverse DNS lookups of crawl logs published by Boa in their transparency report. The bot respects gzip encoding and sends an Accept-Encoding: gzip header. It does not parse JavaScript or execute client-side scripts, focusing exclusively on static HTML, CSS, and structured data (JSON-LD, microdata). Requests are made with a randomized user-agent rotation within the Boa family, but the primary identifier remains the same. The crawler uses a breadth-first traversal strategy and indexes up to 100,000 URLs per domain before applying a random backoff of 10 minutes.
Boa Technologies explicitly states in their robots.txt specification documentation (github.com/boaconstrictor/robots) that boaconstrictor fully honors all Disallow directives, including wildcard patterns and crawl-delay settings. The bot reads the robots.txt file of each domain before every crawl session and caches it with a 24-hour expiry. Evidence from third-party audits (e.g., the Robot Exclusion Compliance Report 2024 by the Internet Policy Institute) confirms that boaconstrictor never violated a Disallow rule in over 2 million tested domains. If a domain returns a 403 or 401 status, the bot immediately aborts further requests to that path.
The primary User-Agent string is BoaConstrictor/2.0 (+https://www.boaconstrictor.io/bot.html), with a fallback string Mozilla/5.0 (compatible; BoaConstrictor/1.8; +https://www.boaconstrictor.io). Behavioral fingerprints include a consistent Accept: text/html,application/xhtml+xml header and a Via header with the value 1.1 boa-proxy. The bot always includes a From: [email protected] header for contact. Network-level detection can rely on the ASN ranges and the fact that crawler requests never include cookies or referrer headers.
Collected content is stored in Boa's Constrictor Cache, an encrypted blob storage system, and used exclusively for training Boa's BoaLLM series of language models and for populating the Constrictor Search Index, a public search engine for developers. The data is not sold to third parties nor used for advertising profiling. Boa publishes a quarterly transparency report detailing all crawled domains, average bandwidth usage, and any complaints received.
Although boaconstrictor is legitimate and well-behaved, it is rate-limited because its concurrent crawling can overwhelm small servers if left unchecked. The recommended threshold for rate limiting is 80 requests per minute per IP, with a 429 response triggering an automatic exponential backoff; this policy is documented in Boa's own rate-limiting guidelines at docs.boaconstrictor.io/rate-limits.
Similar Threats
Free Traffic Analysis
Discover which unwanted bots are being blocked on your site, how often they hit, and where they come from — real data from your own traffic, not guesswork.
🔍 Scan My Site FreePowered by JA4 fingerprinting, honeypot traps & behavioral analysis
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.
Stay up to date with the latest from Boteraser.
We use cookies to improve your experience on our site. By using our site, you consent to cookies.
Manage your cookie preferences below:
Essential cookies enable basic functions and are necessary for the proper function of the website.
CloudFlare provides web performance and security solutions, enhancing site speed and protecting against threats.
Service URL: developers.cloudflare.com (opens in a new window)
These cookies are needed for adding comments on this website.
These cookies are used for managing login functionality on this website.
Statistics cookies collect information anonymously. This information helps us understand how visitors use our website.
Google Analytics is a powerful tool that tracks and analyzes website traffic for informed marketing decisions.
Service URL: policies.google.com (opens in a new window)
You can find more information in our Cookie Policy and Privacy Policy.