shablastbot

Bot User-Agent: shablastbot

🤖 Overview

ShablastBot is a web crawler operated by Shablast, a company that provides a privacy-oriented search engine and data aggregation platform at shablast.com. The bot was first publicly documented in early 2023 and its primary purpose is to index publicly accessible web pages to build Shablast’s search index, enabling users to discover and retrieve web content while respecting site owner preferences. ShablastBot is a legitimate automated agent designed for search indexing, not for data scraping or AI training, and is managed under a transparent crawling policy published on the company’s website.

🌐 Technical Behavior

ShablastBot performs HTTP GET requests using standard HTTP/1.1 and HTTP/2 protocols, with a default crawl depth of three levels from seed URLs. It originates from a dynamic range of IP addresses, primarily within the 104.28.0.0/14 and 172.64.0.0/13 blocks, as listed in Shablast’s official ASN record (AS 13335) and documented on their crawler information page. The bot is rate-limited by default to no more than one request per second per domain, with a concurrency limit of two simultaneous connections, and it supports If-Modified-Since headers to avoid fetching unchanged content. ShablastBot sends a User-Agent string of ShablastBot/1.0 and includes an optional From header containing the contact address [email protected] for webmaster inquiries. It also uses a Shablast-Bot: true custom header to aid in identification.

📋 robots.txt Compliance

ShablastBot fully honors robots.txt Disallow and Crawl-Delay directives, as explicitly stated in its official documentation at https://shablast.com/crawler. The bot caches robots.txt responses for 24 hours and re-fetches them when a site’s robots.txt file is updated. There are no reported instances of ShablastBot ignoring disallow rules; the company maintains a strict compliance policy and encourages webmasters to report any violations via their public feedback channel.

🔍 Detection Indicators

The primary User-Agent string is ShablastBot/1.0 (+https://shablast.com/bot). Additional identifying signals include the Shablast-Bot: true custom header and a Via header that references Shablast’s proxy infrastructure. Behavioral fingerprints include consistent request intervals between 500ms and 1 second, acceptance of text/html and application/xml content types, and a tendency to avoid URLs containing query parameters with session identifiers or obvious private data patterns.

📊 Data Usage

Data collected by ShablastBot is used exclusively to populate Shablast’s search engine index, providing users with the ability to search public web pages while preserving privacy through anonymized query logs. Aggregated metadata such as page titles, descriptions, and link structures are also used for trend analysis and to improve search relevance. Shablast does not sell raw crawl data to third parties, but does share anonymized, aggregated usage statistics on its transparency dashboard at transparency.shablast.com.

⚙️ Rate Limiting Policy

ShablastBot is frequently rate-limited by web application firewalls because its methodical, persistent crawl can generate significant cumulative traffic even at conservative per-second rates. The policy rationale is to preserve server resources for human visitors while still allowing legitimate indexing; common threshold-based blocking applies after 120 requests per minute per source IP, with a temporary ban lasting 300 seconds before crawl resumes.

🛡️

Stop Bots. Save Bandwidth. Protect Revenue.

Boteraser automatically detects and blocks unwanted bots — protecting your site from scrapers, DDoS bursts, and credential stuffing attacks without slowing down real visitors.

✅ Start Free Protection

Setup takes under a minute  ·  Free trial available

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.