Skip to main content

Boteraser | Website and Server Security Solutions

pansophica

Bot User-Agent: pansophica

🤖 Overview

Pansophica is a web crawler operated by Pansophica Inc., a San Francisco-based AI data company that launched in early 2024. It indexes publicly available web content for training large language models and multimodal AI systems, contributing to proprietary datasets and open research. This legitimate bot respects robots.txt and is explicitly rate-limited.

🌐 Technical Behavior

The crawler uses a distributed architecture with IP addresses from cloud providers such as AWS and Google Cloud. It makes about 10 requests per second per IP, using HTTP/1.1 and HTTP/2 with TLS 1.3, and employs HTTP/2 multiplexing for efficiency. It follows recursive link traversal but respects Crawl-Delay directives. The User-Agent string includes a contact email ([email protected]) and a link to its policy page. Pansophica does not execute JavaScript and supports GET and HEAD requests. It checks robots.txt before each crawl.

📋 robots.txt Compliance

Pansophica explicitly honors the Robots Exclusion Standard as documented on its website. It obeys Disallow and Crawl-Delay directives. No violations have been reported in security advisories. The bot also respects X-Robots-Tag HTTP headers.

🔍 Detection Indicators

The primary User-Agent string is Pansophica/1.0 (also "PansophicaBot/1.0"), with version numbers like 1.0.0. Additional headers include From: [email protected] and User-Agent: Pansophica/1.0. Behavioral fingerprints include consistent request intervals and absence of browser-like headers. IP ranges are publicly listed for whitelisting.

📊 Data Usage

Collected data trains generative AI models for natural language understanding and multimodal tasks, and is also used to improve AI safety and reduce bias. Anonymized datasets are shared with academic researchers under open licenses. Data is used internally for model training and fine-tuning, not sold.

⚙️ Rate Limiting Policy

Pansophica is rate-limited due to its distributed traffic potential. Recommended threshold-based blocking at 20 requests per second per IP with temporary bans for high-rate bursts balances server protection with legitimate crawling, as documented in the bot's policy page.

⚠️

Your Site May Be Hemorrhaging Revenue to Bots

Unwanted bots inflate your analytics, drain server resources, and slow down real users. Check if your site is affected — completely free.

Check My Site for Free

Free to start  ·  Cancel anytime

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.