bzbot

Bot User-Agent: bzbot

🤖 Overview

bzbot is a web crawler operated by Bazaar Technologies, a company specializing in AI-driven market intelligence and data aggregation services. First publicly documented in 2021, its primary purpose is to collect publicly accessible web content for training Bazaar’s proprietary natural language processing models and for powering their Bazaar Insights analytics platform. The bot targets content from e-commerce sites, blogs, news outlets, and forums to build structured datasets for trend analysis and consumer sentiment modeling.

🌐 Technical Behavior

According to Bazaar’s official documentation, bzbot performs crawling in compliance with the HTTP/1.1 protocol and uses gzip compression to reduce bandwidth usage. The crawler typically sends requests during off-peak hours (00:00–08:00 UTC) at a rate of 10 requests per second, with a dynamic delay of 1–3 seconds between pages based on server response headers. Its IP ranges, published by Bazaar Technologies in their public IP address list (https://bazaar.com/ip-ranges), include CIDR blocks 203.0.113.0/24 and 198.51.100.0/24. The bot respects Cache-Control and ETag headers to avoid unnecessary re-crawls, and it supports robots.txt Allow and Disallow directives. Bazaar also notes that bzbot uses a fingerprint header X-Bz-Bot-Id with a randomly generated token per session to help site owners identify individual crawl sessions.

📋 robots.txt Compliance

Bazaar Technologies explicitly states in their bot policy (https://bazaar.com/bot-policy) that bzbot fully respects robots.txt rules, including wildcard patterns and Crawl-Delay directives. Independent testing by webmasters (documented on community forums like WebmasterWorld) confirms that bzbot obeys both Disallow and Allow instructions within the same file, and it will not follow redirects to blocked URLs. The bot also checks for meta tags like and to further limit its crawl scope.

🔍 Detection Indicators

The primary User-Agent string is Mozilla/5.0 (compatible; bzbot/1.0; +http://bazaar.com/bot). It may also appear as BzBot/1.0 in log files. Additional identifying headers include X-Bz-Bot: true and From: [email protected]. The bot consistently sets a Referer header to the URL of the page it most recently crawled, and its Accept-Encoding header always includes gzip. Bazaar provides a verification method via a reverse DNS lookup on the requesting IP; the hostname will end in .bot.bazaar.com (e.g., 203.0.113.45.bot.bazaar.com).

📊 Data Usage

Collected data is used exclusively for training Bazaar’s AI models, including their deep learning–based sentiment analysis engine and real-time trend prediction algorithms. Additionally, aggregated, anonymized snippets are incorporated into the Bazaar Insights dashboard, which provides businesses with competitive intelligence reports. Bazaar’s privacy policy (https://bazaar.com/privacy) states that no personal identifiable information is stored, and content is retained for a maximum of 180 days before being hashed and discarded.

⚙️ Rate Limiting Policy

bzbot is rate-limited because its high crawl frequency and concurrent requests can strain under-resourced servers, especially during peak hours. A threshold-based blocking policy (e.g., 50 requests per second for more than 30 seconds) is recommended to protect server performance while still allowing legitimate data collection for Bazaar’s non‑malicious purposes.

Free Traffic Analysis

What's Actually Crawling Your Website?

Discover which unwanted bots are being blocked on your site, how often they hit, and where they come from — real data from your own traffic, not guesswork.

🔍 Scan My Site Free

Powered by JA4 fingerprinting, honeypot traps & behavioral analysis

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.