sphere scout

Bot User-Agent: sphere-scout

🤖 Overview

Sphere Scout is a web crawler operated by Sphere Technologies, a company specializing in large-scale data collection for artificial intelligence model training. First publicly documented in early 2023, the bot's primary purpose is to aggregate publicly available web content to feed into Sphere's proprietary language model, known as SphereLM, as well as for search indexing and knowledge graph enhancement. According to the official Sphere website (sphere.ai), the crawler is designed to be respectful of website policies while enabling high-quality data acquisition.

🌐 Technical Behavior

Sphere Scout employs a distributed crawling architecture using multiple concurrent requests from IP ranges allocated to Amazon Web Services (AWS) and Google Cloud Platform (GCP). The bot typically sends requests at an average rate of one request per second per IP, but may burst up to five requests per second during initial seed discovery. It uses HTTP/1.1 and HTTP/2 protocols, and includes a referrer header pointing to sphere.ai. The crawler follows a breadth-first traversal pattern, prioritizing pages with high PageRank and fresh content. It also parses sitemaps and RSS feeds to discover new URLs. The user agent string includes version numbers, e.g., SphereScout/2.0, and the bot identifies itself via a robot.txt entry for transparency.

📋 robots.txt Compliance

According to the official Sphere documentation (docs.sphere.ai/crawler), Sphere Scout fully honors the robots.txt directives, including Disallow rules with path prefixes and wildcards. It also respects crawl-delay instructions when specified. The bot does not automatically ignore noindex meta tags but does obey nofollow directives on links. There is no documented evidence of intentional violation of robots.txt rules.

🔍 Detection Indicators

The primary User-Agent string is SphereScout/1.0 with variations including Mozilla/5.0 (compatible; SphereScout/2.0; +https://sphere.ai/bot). Behavioral indicators include a consistent request pattern with a fixed delay between page fetches, and the use of Accept-Language header set to en-US. The bot often includes a custom HTTP header X-Sphere-Scout: true for identification. IP addresses are registered under ASN for AWS and GCP, and reverse DNS may resolve to *.sphere-scout.amazonaws.com.

📊 Data Usage

The collected data is used primarily for three purposes: training Sphere's large language models (LLMs) for natural language understanding and generation, building a web-scale knowledge graph for enhanced search results, and improving Sphere's recommendation algorithms. According to a blog post on sphere.ai (dated June 2023), the data is processed and anonymized to remove personally identifiable information before being added to training corpora. No data is sold to third parties.

⚙️ Rate Limiting Policy

Although Sphere Scout is a legitimate and well-behaved crawler, it can generate significant traffic on high-traffic sites due to its distributed nature. Rate limiting is recommended to prevent excessive load on origin servers; a threshold of 10 requests per minute per IP is a common policy to balance data collection with server performance.

🛡️

Stop Bots. Save Bandwidth. Protect Revenue.

Boteraser automatically detects and blocks unwanted bots — protecting your site from scrapers, DDoS bursts, and credential stuffing attacks without slowing down real visitors.

✅ Start Free Protection

Setup takes under a minute  ·  Free trial available

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.