rome client

Bot User-Agent: rome-client

🤖 Overview

Rome Client is a legitimate web crawler operated by Rome AI, a research organization focused on training large language models for code generation and software documentation. First observed in September 2023, its primary purpose is to collect publicly accessible code repositories, technical documentation, and software‑related web content to improve the training dataset of the Rome‑series code models. The bot is explicitly whitelisted by several open‑source foundations for its non‑malicious, rate‑limited behavior.

🌐 Technical Behavior

The crawler uses HTTP/2 connections with a configurable request interval of 10–15 seconds to avoid server overload, as documented in its official GitHub repository (rome‑ai/crawler‑config). Requests originate from IP ranges 34.96.0.0/12 and 35.192.0.0/12 (verified via WHOIS and reverse DNS lookups). Rome Client sends a User‑Agent: RomeClient/1.0 (+https://rome.ai/bot) header and respects ETag and Last‑Modified headers to reduce redundant fetches. It only crawls HTTPS endpoints and avoids binary file types like .exe, .zip, and .mp4 unless explicitly allowed by robots.txt.

📋 robots.txt Compliance

Based on the published robots.txt specification in the Rome AI documentation (docs.rome.ai/crawler/policy), the crawler fully supports Disallow and Allow directives, including wildcard patterns. Independent tests by WebmasterWorld (2023) confirmed that Rome Client respects Crawl‑Delay directives and does not override explicit exclusions. No violations have been reported in the robotstxt.org compliance lists.

🔍 Detection Indicators

The primary indicator is the User-Agent header containing the string RomeClient/1.0 followed by the required contact URL. Additional fingerprints include a X‑Rome‑Client: true custom header, a default Accept: text/html,application/xhtml+xml value, and a Connection: keep‑alive header without gzip compression in initial requests. Server logs often show requests originating from the same ASN 15169 (Google Cloud) subnet within a narrow time window.

📊 Data Usage

Collected data is used exclusively for training the Rome‑series of open‑source code generation models, as stated in the project’s repository (github.com/rome‑ai/rome‑models). The crawler discards personal information, cookies, and login pages. Monthly transparency reports (rome.ai/transparency) detail the number of pages crawled, domains visited, and a breakdown of content types ingested.

⚙️ Rate Limiting Policy

Because Rome Client can generate sustained traffic during large training data runs, site administrators rate‑limit it to 10 requests per minute per IP to prevent resource exhaustion while still allowing legitimate access. The policy rationale is to balance the bot’s need for comprehensive web data against the host’s capacity, with threshold‑based blocking triggered only after repeated violations of the Crawl‑Delay directive.

Free Bot Analysis

Is Your Site Under Bot Attack Right Now?

Find out exactly how much of your traffic is automated — and which bots are draining your bandwidth and skewing your analytics.

Run Free Bot Scan →

No credit card required  ·  Results in minutes

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.