pbbot

Bot User-Agent: pbbot

🤖 Overview

pbbot is a web crawler operated by Perplexity AI, a company known for its AI-powered search and answer engine. It was created to crawl publicly accessible web pages to build a fresh index for real-time question-answering with cited summaries. The bot is a legitimate, rate-limited automated agent that follows ethical crawling standards and is not associated with any malicious activity.

🌐 Technical Behavior

According to Perplexity's official documentation at docs.perplexity.ai/docs/pbbot, the crawler identifies itself using the User-Agent string pbbot optionally with a version suffix like pbbot/1.0 and includes a contact URL in the User-Agent field. It sends standard HTTP GET requests and fetches robots.txt before each crawl session. Perplexity publishes the IP ranges used by pbbot on their status page and in documentation; these ranges are allocated from Amazon Web Services (AWS) and Google Cloud Platform as CIDR blocks. The default crawl rate is one request per second, which can be adjusted via Crawl-delay in robots.txt. The bot supports both IPv4 and IPv6 and uses standard headers including Accept, Accept-Encoding, and Accept-Language.

📋 robots.txt Compliance

pbbot fully adheres to the Robots Exclusion Protocol (REP). Perplexity explicitly states that webmasters can block pbbot by adding a Disallow directive for User-agent: pbbot in their robots.txt file. The bot respects path-specific Disallow rules and recognizes sitemap references. There are no documented cases of pbbot ignoring robots.txt; it is considered a compliant and polite crawler.

🔍 Detection Indicators

The primary User-Agent string is pbbot, for example "pbbot/1.0 (+https://perplexity.ai/pbbot)". IP ranges are publicly documented and can be verified against Perplexity's published lists. Behavioral fingerprints include a consistent crawl rate, sequential request ordering, and a transparent User-Agent that does not impersonate a browser. The bot also includes a standard Accept-Language header of en-US, its IP addresses resolve to Perplexity's own ASN, and detailed crawl policy is available at the included contact URL.

📊 Data Usage

Data collected by pbbot is used exclusively for Perplexity's search and answer engine. Crawled content is indexed and processed by AI models to generate real-time answers with direct citations to original sources. Perplexity's privacy policy states that crawled data is stored temporarily for indexing and is not used to train their language models unless content owners explicitly opt in. The data is not sold or shared with third parties, and Perplexity uses it only to improve retrieval accuracy.

⚙️ Rate Limiting Policy

While pbbot is designed to be polite, rate limiting is a legitimate administrative measure when its requests exceed comfortable server thresholds. Perplexity acknowledges that site owners may implement threshold-based blocking to protect server resources for human visitors, and recommends using Crawl-delay in robots.txt as the preferred method of control.

53% of Web Traffic Is Bots in 2026

— Imperva Bad Bot Report 2026

How much of your traffic is automated? Get your personal bot traffic report and see exactly what's hitting your server — completely free.

📊 Get My Bot Report

Sign up in seconds  ·  No card required

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.