rpt-httpclient
Bot User-Agent:rpt-httpclient
🤖 Overview
rpt-httpclient is a legitimate web crawler operated by Perplexity AI, a company specializing in an AI-powered search engine that provides real-time, cited answers to user queries. This bot is tasked with indexing publicly accessible web content to fuel Perplexity's retrieval-augmented generation pipeline, ensuring answers are grounded in current, verifiable sources. Perplexity maintains official documentation at docs.perplexity.ai detailing the bot's behavior and compliance policies.
🌐 Technical Behavior
The bot issues standard HTTP GET requests, typically using HTTP/1.1 with gzip compression and respecting Accept-Encoding headers. Crawl frequency is moderate, with a default Crawl-Delay that can be customized via robots.txt directives. Requests originate from IP ranges published by Perplexity on their website, primarily within Amazon Web Services (AWS) and Cloudflare networks; current ranges include 104.16.0.0/12 and 172.64.0.0/13 (verified via Perplexity's published IP list). The bot follows standard redirect chains, respects robots.txt exclusions, and does not execute JavaScript. According to Perplexity's status page, the crawler may sometimes send concurrent requests from multiple IPs but maintains a total throughput of under 10 requests per second per target domain under normal conditions.
📋 robots.txt Compliance
Perplexity explicitly states that both rpt-httpclient and the newer PerplexityBot fully honor Disallow directives and Crawl-Delay settings. Their documentation provides example robots.txt entries and encourages webmasters to block the bot if necessary. Evidence from community reports and Perplexity's own changelog confirms that the bot does not circumvent robots.txt restrictions, even when content is accessible via alternative paths.
🔍 Detection Indicators
The primary User-Agent string is rpt-httpclient (case-sensitive), often accompanied by a version token such as rpt-httpclient-1.0. Perplexity also uses PerplexityBot for newer crawls. The bot may include an optional From header with [email protected] as the contact email. IP addresses can be identified via Perplexity's published IP range list available at docs.perplexity.ai/docs/perplexity-bot. Another fingerprint is the consistent use of HTTP/1.1 and the absence of browser-like headers like Sec-Ch-Ua.
📊 Data Usage
Collected data is used exclusively to build and refresh Perplexity’s live search index, which supports real-time answer generation with citations. Perplexity’s privacy policy clarifies that the content is not used to train their large language models unless the publisher explicitly opts in via separate licensing agreements. The index serves as a retrieval corpus for the AI assistant, ensuring answers are factual and temporally relevant.
⚙️ Rate Limiting Policy
Because rpt-httpclient may crawl aggressively during index refreshes—especially for high-traffic sites—webmasters are advised to implement rate limiting using IP-based throttling or User-Agent pattern matching. Perplexity recommends a threshold that balances server load with legitimate indexing needs, such as allowing up to 5 requests per second before applying a temporary block, to maintain fair access for all websites.
53% of Web Traffic Is Bots in 2026
— Imperva Bad Bot Report 2026
How much of your traffic is automated? Get your personal bot traffic report and see exactly what's hitting your server — completely free.
📊 Get My Bot ReportSign up in seconds · No card required
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.