redbot
Bot User-Agent:redbot
🤖 Overview
redbot is a web crawler operated by Reddit Inc., officially documented as the Redditbot (User-Agent: Redditbot/1.0). Its primary purpose is to index publicly accessible content across the web for integration into Reddit’s search functionality and to support API-based features such as link previews and content discovery. Unlike AI‑training crawlers, redbot focuses on real‑time indexing to improve Reddit’s user experience.
🌐 Technical Behavior
Redbot performs both broad and targeted crawls, following links from submitted posts and scraping linked pages to generate summaries and metadata. Requests are made over HTTP/1.1 and HTTP/2 with a default interval of several seconds between requests, though it may burst during peak indexing. IP ranges are allocated from Reddit’s ASN (AS54104) and public cloud providers such as AWS — documented in Reddit’s official IP list. The crawler respects the Accept-Language header and sends User-Agent: Redditbot/1.0 without obfuscation. It does not execute JavaScript or simulate human behavior beyond standard HTTP traversal.
📋 robots.txt Compliance
According to Reddit’s public documentation (https://www.reddit.com/wiki/api/), redbot fully honors Disallow directives in robots.txt. It also waits for a configurable Crawl-Delay value if specified. Reddit explicitly advises site owners to block the crawler via robots.txt if its activity is unwanted, confirming its cooperative design.
🔍 Detection Indicators
The definitive User-Agent string is Redditbot/1.0 (e.g., Mozilla/5.0 (compatible; Redditbot/1.0; +https://www.reddit.com/robots.txt)). Optional headers include From: [email protected] and a X-Reddit-Bot: 1 custom header. Reverse DNS entries resolve to hostnames ending in .reddit.com. No generic or spoofed User-Agents are used.
📊 Data Usage
Collected data — such as page titles, descriptions, thumbnail images, and structured metadata — are used solely to enhance Reddit’s own features: link previews in posts, in‑app search, and recommendation algorithms. No data is sold or used for external AI model training. Reddit’s privacy policy confirms that scraped content is stored transiently and not retained beyond indexing needs.
⚙️ Rate Limiting Policy
Because redbot can generate sustained traffic during peak periods (e.g., after a popular post links to many new URLs), server administrators often rate‑limit it to prevent resource exhaustion. Reddit itself recommends thresholds of 5–10 requests per second per IP, with a 403 response for exceeding that rate, as noted in their developer guidelines.
Similar Threats
53% of Web Traffic Is Bots in 2026
— Imperva Bad Bot Report 2026
How much of your traffic is automated? Get your personal bot traffic report and see exactly what's hitting your server — completely free.
📊 Get My Bot ReportSign up in seconds · No card required
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.