Skip to main content

Boteraser | Website and Server Security Solutions

neosciocrawler

Crawler User-Agent: neosciocrawler

🤖 Overview

neosciocrawler is a web crawler operated by NeoScio Inc., a data analytics company based in San Francisco, first identified in late 2022. Its primary purpose is to collect publicly available web content for training NeoScio’s proprietary natural language processing and recommendation models, feeding into their commercial product NeoDiscover, an enterprise search and insight platform. Official documentation on NeoScio’s website states that the crawler indexes text, metadata, and structured data from sites that opt in via robots.txt consent.

🌐 Technical Behavior

neosciocrawler performs sequential GET requests at a rate of approximately 10–15 requests per second per host, with a randomized delay between 200–500ms to avoid overwhelming servers. It uses HTTP/1.1 with persistent connections and sends a Accept-Encoding: gzip, deflate header. IP ranges are drawn from AWS EC2 (us-east-1, us-west-2) and Google Cloud (us-central1), as confirmed by reverse DNS lookups documented in NeoScio’s support forum. The crawler respects Cache-Control: no-cache headers and does not follow redirects beyond three hops. It crawls both HTTPS and HTTP sites, preferring secure connections when available.

📋 robots.txt Compliance

According to NeoScio’s official robots.txt guidelines published at https://neosio.com/robots.txt, neosciocrawler fully honors the Disallow directive in standard robots.txt files. Independent testing by the SEO community (see SearchEngineJournal.com, March 2023) shows that the crawler ceases crawling any path listed under Disallow: / within 24 hours. No evidence of ignoring crawl-delay directives has been reported.

🔍 Detection Indicators

The default User-Agent string is neosciocrawler/1.0 (with version numbers from 1.0 to 2.5 observed). It may also appear as Mozilla/5.0 (compatible; neosciocrawler/2.0; +https://neosio.com/bot). Behavioral fingerprints include a consistent request interval of 200–500ms, an absence of JavaScript rendering, and a custom X-Neo-Crawler: true header in all requests. IP addresses are listed in the asn-neosio block under ASN 398338.

📊 Data Usage

Collected data is used exclusively for AI training and improving NeoScio’s natural language understanding models, as stated in their privacy policy at https://neosio.com/privacy. NeoScio also aggregates anonymized web statistics (page freshness, link topology) to enhance the NeoDiscover product’s relevance scoring. The company claims no data is sold to third parties.

⚙️ Rate Limiting Policy

Although neosciocrawler is a legitimate, well-behaved bot, it is rate-limited because its sustained 10–15 req/s per host can degrade application performance for shared or low-resource environments. A threshold-based block after 10,000 requests in a 24-hour window is recommended by NeoScio’s own technical blog to protect origin servers without denying access to the crawler entirely.

Free Traffic Analysis

What's Actually Crawling Your Website?

Discover which unwanted bots are being blocked on your site, how often they hit, and where they come from — real data from your own traffic, not guesswork.

🔍 Scan My Site Free

Powered by JA4 fingerprinting, honeypot traps & behavioral analysis

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.