Skip to main content

Boteraser | Website and Server Security Solutions

scich

Bot User-Agent: scich

🤖 Overview

Scich is a web crawler operated by Scich Inc., a data analytics company specializing in scientific literature indexing and AI-assisted research discovery. First publicly documented in 2022, the bot systematically collects publicly available academic content – including preprint repositories, publisher websites, and institutional repositories – to feed the Scich platform, which provides citation analysis, trend detection, and natural-language search over millions of papers. According to the official Scich documentation (https://scich.com/robots), the crawler is designed to respect publisher terms and is used exclusively for non‑commercial research enrichment.

🌐 Technical Behavior

Scich performs sequential, single‑threaded crawling with a default request delay of 10 seconds between pages, as backed by its published crawl policy (https://scich.com/crawl-policy). It issues HTTP GET requests with the Accept‑Encoding: gzip header and advertises its identity via the User‑Agent string: Mozilla/5.0 (compatible; ScichBot/1.0; +https://scich.com/bot). The bot primarily targets HTML pages, PDF files, and XML feeds (e.g., sitemaps), and it follows HTTP redirects but does not crawl JavaScript‑rendered content. IP ranges allocated to Scich are listed in the AS‑SCICH autonomous system (ASN 393939) and are fully documented in the Scich IP Range Registry (https://scich.com/ip-ranges). The crawler supports both HTTP/1.1 and HTTP/2 protocols and respects the Cache‑Control header to reduce server load.

📋 robots.txt Compliance

Scich strictly honors robots.txt Disallow directives and also respects the Crawl‑Delay directive if specified. Multiple publisher tests (e.g., from Elsevier and arXiv) have confirmed that the bot does not access disallowed paths and waits for the defined delay before each request. The official Scich documentation explicitly states that the crawler will never bypass robots.txt or attempt hidden directories.

🔍 Detection Indicators

The primary identifier is the User‑Agent string: Mozilla/5.0 (compatible; ScichBot/1.0; +https://scich.com/bot). Behavioral fingerprints include a consistent 10‑second inter‑request interval and a strong preference for academic document URLs containing .pdf, doi.org, or arxiv.org paths. The bot also sends a custom X‑Scich‑Client: v1.0 header in every request, which can be used for log‑based identification.

📊 Data Usage

Collected content is ingested into the Scich Knowledge Graph, where it is parsed for citations, metadata, and full‑text passages. The processed data supports the Scich search engine (scich.com/search), citation recommendation systems, and research trend visualizations. No raw content is sold or redistributed; instead, aggregated statistics and summary embeddings are used for AI‑powered literature discovery.

⚙️ Rate Limiting Policy

Although Scich is legitimate, its aggressive crawl pattern can consume significant server resources on high‑traffic sites, especially when it encounters large PDF archives. System administrators are advised to apply moderate rate‑limiting (e.g., 10 requests per minute) to protect application performance while still allowing the bot to gather necessary data for the academic research ecosystem.

Free Bot Analysis

Is Your Site Under Bot Attack Right Now?

Find out exactly how much of your traffic is automated — and which bots are draining your bandwidth and skewing your analytics.

Run Free Bot Scan →

No credit card required  ·  Results in minutes

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.