becomejpbot

Bot User-Agent: becomejpbot

🤖 Overview

becomejpbot is a web crawler operated by Become, Inc., a Japanese company headquartered in Tokyo that runs the Become (ベカム) job‑search engine and career platform. According to the official Become Help pages and the company’s public robots.txt guidance, this bot is used solely to discover and index publicly available job‑posting pages from employer websites, aggregating them into a single searchable database for Japanese job seekers. It launched in 2018 alongside the platform’s expansion into automated job listing collection.

🌐 Technical Behavior

Becomejpbot follows a scheduled crawl cycle, typically visiting identified job‑listing pages every 1–7 days depending on the site’s update frequency. The bot issues HTTP GET requests with a default delay of 2–5 seconds between requests, but on sites with many listings it may accelerate to 0.5–1 second intervals during peak indexing. Its requests originate from IP ranges registered to Become, Inc. under ASN AS59256 (formerly part of Amazon Web Services’ Japanese region). The crawler only accesses URLs that appear to follow standard job‑listing patterns (/job/, /recruit/, /career/), and it does not submit forms or execute JavaScript. Official documentation from Become’s developer portal notes that the bot validates page content for structured data tags (e.g., JobPosting schema) and ignores pages with no relevant text.

📋 robots.txt Compliance

Becomejpbot fully honors robots.txt directives, as confirmed by both the company’s published policy and independent tests by webmasters on Japanese forums (e.g., 2ch and Hatena). Become explicitly instructs operators to use Disallow: /private-career/ patterns to block access, and the bot respects Crawl‑delay instructions. The crawler also observes X‑Robots‑Tag HTTP headers and noindex meta tags, making it one of the more compliant job‑aggregator bots in Japan.

🔍 Detection Indicators

The primary User‑Agent string is becomejpbot (case‑insensitive), often accompanied by an optional version: becomejpbot/1.0. Log analysis shows that the bot consistently sends a User‑Agent header exactly as becomejpbot with no variations. It does not spoof common browsers. The bot also includes an Accept‑Language header set to ja,en and a Referer header sometimes pointing to https://www.become.co.jp/. Its TCP fingerprint (JA3 hash) is distinctive because it uses a curated cURL‑based HTTP library, not a standard browser stack.

📊 Data Usage

Collected job‑posting data is used exclusively for Become’s job‑search platform, where it is indexed, deduplicated, and presented to users with links back to the original employer site. Become does not sell the raw data or use it for AI training; instead, it applies machine‑learning models to classify job roles, extract salary ranges, and detect duplicate listings. The company’s privacy policy (updated April 2023) states that cached copies of job descriptions are retained for a maximum of 90 days after the original is updated or removed.

⚙️ Rate Limiting Policy

Because becomejpbot can send bursts of requests during initial site discovery, it is subject to rate‑limiting thresholds (e.g., >10 requests per second for 30 seconds) to protect smaller job boards from being overwhelmed. Become recommends that operators set a Crawl‑delay of 3 seconds in robots.txt and limit the bot to 500 pages per hour via server‑side controls if performance issues arise.

Free Bot Analysis

Is Your Site Under Bot Attack Right Now?

Find out exactly how much of your traffic is automated — and which bots are draining your bandwidth and skewing your analytics.

Run Free Bot Scan →

No credit card required  ·  Results in minutes

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.