redbot
redbot is a web crawler operated by Reddit Inc., officially documented as the Redditbot (User-Agent: Redditbot/1.0). Its primary purpose is to index publicly accessible content across the web for integration into Reddit’s search functionality and to support API-based features such as link previews and content discovery. Unlike AI‑training crawlers, redbot focuses on real‑time indexing to improve Reddit’s user experience.
Redbot performs both broad and targeted crawls, following links from submitted posts and scraping linked pages to generate summaries and metadata. Requests are made over HTTP/1.1 and HTTP/2 with a default interval of several seconds between requests, though it may burst during peak indexing. IP ranges are allocated from Reddit’s ASN (AS54104) and public cloud providers such as AWS — documented in Reddit’s official IP list. The crawler respects the Accept-Language header and sends User-Agent: Redditbot/1.0 without obfuscation. It does not execute JavaScript or simulate human behavior beyond standard HTTP traversal.
According to Reddit’s public documentation (https://www.reddit.com/wiki/api/), redbot fully honors Disallow directives in robots.txt. It also waits for a configurable Crawl-Delay value if specified. Reddit explicitly advises site owners to block the crawler via robots.txt if its activity is unwanted, confirming its cooperative design.
The definitive User-Agent string is Redditbot/1.0 (e.g., Mozilla/5.0 (compatible; Redditbot/1.0; +https://www.reddit.com/robots.txt)). Optional headers include From: [email protected] and a X-Reddit-Bot: 1 custom header. Reverse DNS entries resolve to hostnames ending in .reddit.com. No generic or spoofed User-Agents are used.
Collected data — such as page titles, descriptions, thumbnail images, and structured metadata — are used solely to enhance Reddit’s own features: link previews in posts, in‑app search, and recommendation algorithms. No data is sold or used for external AI model training. Reddit’s privacy policy confirms that scraped content is stored transiently and not retained beyond indexing needs.
Because redbot can generate sustained traffic during peak periods (e.g., after a popular post links to many new URLs), server administrators often rate‑limit it to prevent resource exhaustion. Reddit itself recommends thresholds of 5–10 requests per second per IP, with a 403 response for exceeding that rate, as noted in their developer guidelines.
Similar Threats
Free Traffic Analysis
Discover which unwanted bots are being blocked on your site, how often they hit, and where they come from — real data from your own traffic, not guesswork.
🔍 Scan My Site FreePowered by JA4 fingerprinting, honeypot traps & behavioral analysis
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.
Stay up to date with the latest from Boteraser.
We use cookies to improve your experience on our site. By using our site, you consent to cookies.
Manage your cookie preferences below:
Essential cookies enable basic functions and are necessary for the proper function of the website.
CloudFlare provides web performance and security solutions, enhancing site speed and protecting against threats.
Service URL: developers.cloudflare.com (opens in a new window)
These cookies are needed for adding comments on this website.
These cookies are used for managing login functionality on this website.
Statistics cookies collect information anonymously. This information helps us understand how visitors use our website.
Google Analytics is a powerful tool that tracks and analyzes website traffic for informed marketing decisions.
Service URL: policies.google.com (opens in a new window)
You can find more information in our Cookie Policy and Privacy Policy.