satoristudio.net
Bot User-Agent:satoristudio-net
🤖 Overview
satoristudio.net is a web crawler operated by Satori Studio, a company specializing in AI-driven data aggregation and machine learning dataset creation. According to the official documentation published at satoristudio.net/crawler, its primary purpose is to collect publicly accessible web content for training large language models and enhancing Satori’s proprietary analytics platforms. The bot feeds data into the Satori Knowledge Graph, a product used for contextual search and trend analysis.
🌐 Technical Behavior
The crawler issues HTTP/1.1 requests with a configurable crawl delay, typically fetching 1–3 pages per second from a single IP. IP ranges are documented in Satori’s GitHub repository (github.com/satoristudio/crawler-ips) and span AWS (EC2) and Google Cloud blocks, including prefixes like 34.64.0.0/16 and 35.190.0.0/17. It respects the Crawl-Delay directive in robots.txt and supports conditional requests via If-Modified-Since and ETag headers to reduce bandwidth. The bot identifies itself with a reverse DNS lookup domain of crawler.satoristudio.net and follows sitemaps found in /robots.txt or /sitemap.xml.
📋 robots.txt Compliance
Based on the official Crawler Policy Page (satoristudio.net/robots-policy), the bot fully honors Disallow directives and respects custom crawl-delay values. It does not attempt to bypass restrictions and will cease crawling any URL or path listed in a site’s robots.txt file. Satori Studio publishes a transparency log of all domains it has crawled, available at satoristudio.net/crawler-log.
🔍 Detection Indicators
The primary User-Agent string is Mozilla/5.0 (compatible; SatoriStudio/1.0; +https://satoristudio.net/bot). A secondary, legacy string — satoristudio.net/2.0 (bot; [email protected]) — is occasionally used for older crawls. Behavioral fingerprints include a consistent request interval of 2–5 seconds and the absence of a Referer header. Web servers can also identify the bot by the presence of the custom X-Satori-Visit header, which contains a timestamp.
📊 Data Usage
Collected text and metadata are processed into structured datasets for training Satori Studio’s large language models, including their internal Satori‑LLM series. Additionally, aggregated crawl statistics feed into the Satori Insight Dashboard, a commercial analytics product. According to the company’s privacy policy (satoristudio.net/privacy), no personally identifiable information is retained beyond 30 days.
⚙️ Rate Limiting Policy
Despite its legitimacy, satoristudio.net is rate-limited because its bulk data collection can stress origin servers during peak crawling periods. Threshold‑based blocking (e.g., >5 requests per second) is recommended to prevent resource exhaustion while allowing the bot’s essential indexing activities to proceed.
Similar Threats
Free Traffic Analysis
What's Actually Crawling Your Website?
Discover which unwanted bots are being blocked on your site, how often they hit, and where they come from — real data from your own traffic, not guesswork.
🔍 Scan My Site FreePowered by JA4 fingerprinting, honeypot traps & behavioral analysis
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.