becomebot
Bot User-Agent:becomebot
🤖 Overview
BecomeBot is a web crawler operated by Become, Inc., a company that runs the shopping comparison engine Become.com and its international variants such as Become.co.uk and Become.co.jp. Its primary purpose is to index product listings, pricing data, merchant inventory, and consumer reviews from e-commerce websites to feed the Become shopping platform, which aggregates product information for price comparison and online shopping. The bot is a legitimate commercial agent used exclusively for deal aggregation, not for AI training or search indexing.
🌐 Technical Behavior
BecomeBot performs HTTP/1.1 GET requests on public product URLs, typically following links from sitemaps and category pages. It respects a default crawl delay of 10 seconds between consecutive requests to the same host, as stated in official documentation on Become’s developer help page. The bot originates from a small set of IPv4 addresses predominantly within the 66.155.9.0/24 and 66.155.10.0/24 ranges, allocated by ARIN to Become, Inc., confirmed via reverse DNS lookups showing hostnames like become.bot.prod.become.com. It does not render JavaScript or execute client-side scripts, focusing exclusively on raw HTML and structured data such as schema.org microdata and JSON-LD pricing schemas. The crawler identifies itself via the User-Agent string typically reading "Mozilla/5.0 (compatible; BecomeBot/3.0; +http://www.become.com/site_owners.html)", though older versions (e.g., 2.0) are also observed in the wild.
📋 robots.txt Compliance
According to Become’s site owner guide (available at http://www.become.com/site_owners/bot.html), BecomeBot honors Disallow directives specified in the robots.txt file, and site owners are encouraged to use standard rules to block specific directories. The bot also respects the Crawl-Delay directive if set in robots.txt, allowing webmasters to further throttle its request rate. There is no evidence of the bot ignoring robots.txt or bypassing access controls.
🔍 Detection Indicators
The primary detection fingerprint is the User-Agent string: "Mozilla/5.0 (compatible; BecomeBot/3.0; +http://www.become.com/site_owners.html)". Additional versions include "BecomeBot/2.0" and "BecomeBot/1.0". The bot also sends a referer header often set to the main Become.com domain and typically does not include an Accept-Encoding header for gzip compression. Reverse DNS lookups connecting back to *.prod.become.com help confirm its identity. No known CVE identifiers or security advisories reference BecomeBot as a vector.
📊 Data Usage
Collected data—including product titles, prices, availability, descriptions, merchant names, and customer ratings—is used exclusively to populate the Become.com comparison shopping engine. This data is not used for AI training, language model development, or any generative purposes. Become, Inc. states on its site owner page that the data is only published as searchable listings and may be cached for up to 48 hours to reflect real-time pricing changes.
⚙️ Rate Limiting Policy
Rate limiting is recommended because the bot may send bursts of requests when initially indexing a large catalog, even with its built-in delay. A threshold-based blocking mechanism (e.g., >50 requests per minute) ensures fair resource usage without permanently blocking the crawler, which is necessary for legitimate price comparison functionality.
Similar Threats
53% of Web Traffic Is Bots in 2026
— Imperva Bad Bot Report 2026
How much of your traffic is automated? Get your personal bot traffic report and see exactly what's hitting your server — completely free.
📊 Get My Bot ReportSign up in seconds · No card required
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.