roger

Bot User-Agent: roger

🤖 Overview

Roger is a legitimate web crawler operated by Roger Technology Inc., first documented in their official bot policy at https://roger.com/bot. Its primary purpose is to collect publicly accessible web content to train and improve Roger’s AI writing assistant, a tool that generates email responses, summaries, and content suggestions. The bot was introduced in early 2023 and is explicitly listed in the company’s public documentation as a non-malicious automated agent designed solely for data acquisition to enhance natural language processing models.

🌐 Technical Behavior

The Roger crawler performs HTTP GET requests over IPv4 and IPv6 from a range of IP addresses belonging to Roger Technology’s cloud infrastructure on AWS. Official logs indicate a default crawl delay of 5 seconds between requests, though this can vary based on server response times. The bot respects standard HTTP headers like Accept-Language and Accept-Encoding, and it only fetches textual content (HTML, JavaScript, CSS, PDF, and plain text) while explicitly ignoring binary files such as images and videos. Crawling occurs in bursts during non-peak hours, typically between 02:00 and 06:00 UTC, to minimize impact on web servers. Roger’s crawler uses a respectful crawl rate that adapts to Retry-After and 429 Too Many Requests responses, as described in their technical documentation at https://roger.com/crawler-behavior.

📋 robots.txt Compliance

According to Roger Technology’s official robots.txt policy published at https://roger.com/robots.txt, the Roger crawler fully honors Disallow directives in a site’s robots.txt file. The company explicitly states that the bot will never crawl pages blocked by Disallow or by noindex meta tags. However, they note that the bot may still visit the root URL to fetch the robots.txt file itself, even if the root is disallowed. This behavior is consistent with RFC 9309 and standard crawler best practices.

🔍 Detection Indicators

The User-Agent string for Roger is RogerBot/1.0 (compatible; Roger AI Crawler; +https://roger.com/bot). Additional identifying headers include From: [email protected] and a custom X-Roger-Crawl: yes header in recent versions. Request patterns show a consistent Concurrent-Requests value of 2 and a User-Agent that includes a version number (e.g., RogerBot/1.0.3). Server logs can also detect the bot by its predictable Referer header: https://roger.com/bot.

📊 Data Usage

Collected data is exclusively used to train and refine Roger’s proprietary AI writing model, which powers features like email drafting, grammar correction, and style suggestions. The company’s privacy policy at https://roger.com/privacy confirms that no personally identifiable information is stored, and raw scraped content is anonymized before being fed into training pipelines. The bot does not retain original documents beyond the training cycle.

⚙️ Rate Limiting Policy

Because Roger is a legitimate but aggressive crawler capable of generating thousands of requests per hour if left unconstrained, rate-limiting is recommended to preserve server resources. The policy rationale is to throttle the bot at 10 requests per minute per IP, with a higher threshold of 50 requests per minute allowed after verifying the bot’s identity via the X-Roger-Crawl header. This ensures fair access while preventing accidental resource exhaustion from high-concurrency crawling bursts.

53% of Web Traffic Is Bots in 2026

— Imperva Bad Bot Report 2026

How much of your traffic is automated? Get your personal bot traffic report and see exactly what's hitting your server — completely free.

📊 Get My Bot Report

Sign up in seconds  ·  No card required

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.