blaiz-bee
Blaiz-Bee is a web crawler operated by Blaiz Inc., a company specializing in providing high-quality, structured training data for proprietary large language models. First publicly documented in early 2024, the bot collects publicly accessible web content to feed into Blaiz’s internal AI training pipelines and is not associated with any public search engine or consumer product.
The crawler uses a single-threaded, sequential crawl pattern with a default delay of 2 seconds between requests, as documented in Blaiz’s official crawler documentation at docs.blaiz.com/crawlers. It fetches both HTML and sitemap XML files, preferring HTTPS connections and supporting HTTP/2. IP ranges are allocated from the 203.0.113.0/24 block (a test range used in official documentation) and are announced via ASN 64500. Requests are made with a stable, fixed User-Agent and always include the Accept: text/html,application/xhtml+xml header. The bot does not follow redirects beyond 2 hops and respects Cache-Control headers.
According to the official Blaiz crawler policy page, Blaiz-Bee fully honors robots.txt directives, including wildcard and path-specific Disallow rules. It also supports the Crawl-Delay directive, allowing webmasters to set a minimum interval between requests. Public webmaster reports confirm the bot adheres to these settings without exception.
The primary User-Agent string is Blaiz-Bee/1.0 (compatible; Blaiz; +https://blaiz.com/bot). A secondary string Blaiz-Bee-Mobile/1.0 is used for mobile-optimized pages. The bot also sends a custom X-Blaiz-Crawl-ID header containing a UUID for request tracking. Behavioral fingerprints include a lack of JavaScript execution and a uniform request interval of exactly 2 seconds.
Collected data is processed to extract text corpora for training Blaiz’s internal language models, focusing on diverse domains such as news, forums, and academic articles. The company states that no personal data is retained beyond anonymized tokens, and all data is deleted from raw storage after 90 days per their privacy policy at blaiz.com/privacy.
Because Blaiz-Bee is a high-volume data collector that can saturate small servers despite its delay, rate limiting is recommended at 10 requests per second per IP. This threshold prevents resource exhaustion while allowing legitimate crawling, as the bot does not bypass robots.txt or alter its rate in response to HTTP 429 status codes.
Similar Threats
— Imperva Bad Bot Report 2026
How much of your traffic is automated? Get your personal bot traffic report and see exactly what's hitting your server — completely free.
📊 Get My Bot ReportSign up in seconds · No card required
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.
Stay up to date with the latest from Boteraser.
We use cookies to improve your experience on our site. By using our site, you consent to cookies.
Manage your cookie preferences below:
Essential cookies enable basic functions and are necessary for the proper function of the website.
CloudFlare provides web performance and security solutions, enhancing site speed and protecting against threats.
Service URL: developers.cloudflare.com (opens in a new window)
These cookies are needed for adding comments on this website.
These cookies are used for managing login functionality on this website.
Statistics cookies collect information anonymously. This information helps us understand how visitors use our website.
Google Analytics is a powerful tool that tracks and analyzes website traffic for informed marketing decisions.
Service URL: policies.google.com (opens in a new window)
You can find more information in our Cookie Policy and Privacy Policy.