Panscient

Bot User-Agent: panscient

🤖 Overview

Panscient is a web crawler operated by Panscient Inc., a company headquartered in Seattle, Washington, that specializes in large-scale web data collection for artificial intelligence training datasets. First publicly documented in 2023, the bot systematically indexes publicly accessible web content including text, images, and structured data to feed into proprietary AI model training pipelines. Unlike general-purpose search engine crawlers, Panscient focuses on building high-quality, diverse datasets for deep learning applications.

🌐 Technical Behavior

The crawler employs a distributed architecture using IP addresses from cloud providers such as AWS and Google Cloud, with a typical request rate of 1–2 requests per second per IP. It follows a breadth-first traversal strategy and supports both HTTP/1.1 and HTTPS, sending requests with a configured Accept header that prefers HTML and plain text. Panscient respects standard HTTP cache-control headers and preserves existing session cookies when present. Analysis of its crawl logs reveals it often requests robots.txt before any page fetch, and its requests include a User-Agent string that is easily identifiable.

📋 robots.txt Compliance

Panscient explicitly states on its official website (panscient.com/robots) that it fully honors robots.txt Disallow directives. Multiple webmaster reports confirm that after adding Disallow rules for the bot’s User-Agent, the crawler ceased accessing those paths within 24 hours. It also respects the Crawl-Delay directive where specified.

🔍 Detection Indicators

The primary User-Agent string is "Panscient/1.0" (or "Panscient/2.0" for newer versions). No additional custom HTTP headers are sent. The bot does not rotate its User-Agent per request, making it straightforward to identify via server logs. It originates from IPv4 ranges registered to Amazon and Google, with reverse DNS names often containing "panscient".

📊 Data Usage

Collected data is used to train large language models, multimodal AI systems, and other commercial machine learning products offered by Panscient’s clients. The company publishes a privacy policy on its site explaining that content is cached, processed, and stored in secure cloud environments for model training and fine-tuning. No personal identifiable information (PII) is intentionally harvested, though the crawler cannot filter PII from public pages.

⚙️ Rate Limiting Policy

Because Panscient can generate sustained traffic of up to several hundred requests per minute from a single IP range, web application owners are advised to rate-limit requests at a threshold of 20 requests per second per IP. This policy protects server resources while still allowing legitimate crawling. Panscient itself recommends a minimum rate limit of 5 requests per second to avoid overloading small websites.

Free Bot Analysis

Is Your Site Under Bot Attack Right Now?

Find out exactly how much of your traffic is automated — and which bots are draining your bandwidth and skewing your analytics.

Run Free Bot Scan →

No credit card required  ·  Results in minutes

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.