spider terranautic net

Crawler User-Agent: spider-terranautic-net

🤖 Overview

spider terranautic net is a web crawler operated by Terranautic Inc., a data analytics firm headquartered in Austin, Texas, as documented on their official bot information page at https://terranautic.net/bot. The bot is designed to collect publicly accessible web content for Terranautic’s competitive intelligence and market trend analysis platform, which serves enterprise clients in retail, finance, and technology sectors. The company explicitly states that the crawler does not scrape personal or login-protected data and complies with all applicable data privacy regulations including the GDPR and CCPA.

🌐 Technical Behavior

The crawler employs a distributed architecture with IP ranges sourced from AWS (prefixes 35.180.0.0/16, 34.96.0.0/16) and Google Cloud (35.191.0.0/16), as published in Terranautic’s IP whitelist at https://terranautic.net/ips.txt. Request frequency defaults to 2–3 requests per second per IP, but administrators can specify a crawl-delay directive in robots.txt to reduce this rate. The bot supports HTTP/1.1 and HTTP/2 protocols, and uses a headless Chromium browser (version 112+) to render JavaScript-heavy pages, as noted in their technical documentation. It recursively follows links up to a depth of 5 by default, and respects both the X-Robots-Tag header and meta robots tags. The crawler sends an Accept-Encoding: gzip, deflate header and includes a unique request ID in the X-Request-ID HTTP header for logging purposes, according to a 2023 whitepaper published by the company.

📋 robots.txt Compliance

Terranautic publicly states on their bot information page that spider terranautic net fully honors robots.txt Disallow directives and supports the Crawl-Delay parameter. Independent testing by BotCheck.com in July 2024 confirmed that the bot reads the file at the start of each crawl session, caches it for 24 hours, and refuses to visit any disallowed paths, including /admin or /private directories. The company also provides a contact email ([email protected]) for site owners to request custom exclusions.

🔍 Detection Indicators

The primary User-Agent string is ‘Mozilla/5.0 (compatible; spider; +http://terranautic.net/bot)’. Alternate strings include ‘TerranauticBot/1.0’ and ‘spider/1.0’, as listed in the official documentation. The bot always sets a custom HTTP header ‘X-Terranautic-Crawler: true’ for easy identification, and sends a ‘From’ header containing the email address ([email protected]). Behavioral fingerprints include sequential request patterns and a fixed accept-language of ‘en-US,en;q=0.9’.

📊 Data Usage

Collected data is aggregated into Terranautic’s proprietary database to produce real-time dashboards for price monitoring, product availability, and brand sentiment analysis, as described in their product documentation at https://terranautic.net/product. The company explicitly states that no data is used for AI model training or search indexing; all data is anonymized and encrypted (AES-256) at rest. Retention periods are limited to 90 days, after which raw data is deleted and only aggregated statistics are kept, in compliance with their published data policy.

⚙️ Rate Limiting Policy

Rate limiting is applied because the crawler’s distributed IP pool can generate up to 1,200 requests per minute across multiple sources, which may overwhelm under‑provisioned servers; Terranautic recommends a threshold‑based block at 500 requests per minute per IP with a 10‑minute cooldown, as outlined in their rate limiting guide. This policy balances the bot’s legitimate data collection needs with the site owner’s performance requirements, and is not a measure to block malicious activity.

Free Traffic Analysis

What's Actually Crawling Your Website?

Discover which unwanted bots are being blocked on your site, how often they hit, and where they come from — real data from your own traffic, not guesswork.

🔍 Scan My Site Free

Powered by JA4 fingerprinting, honeypot traps & behavioral analysis

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.