supybot
Supybot is a web crawler operated by SuPy Inc., a company specializing in large-scale language model training data collection. First documented in public logs around 2021, the crawler's primary purpose is to gather publicly accessible web content for improving SuPy's proprietary AI models and for use in its internal search index. The bot is also utilized by third-party researchers under license via SuPy's data marketplace platform.
The Supybot crawler employs a multi-threaded, asynchronous request pattern using Python's aiohttp library. It typically dispatches between 10 to 20 requests per second per IP, with bursts reaching up to 50 requests per second during peak indexing cycles. The bot primarily uses HTTP/1.1 and HTTP/2 with keep-alive connections. IP ranges are drawn from ASN AS394256, which includes addresses like 45.33.32.0/19 and 2600:3c00::/32 (dual-stack IPv4/IPv6). Crawl depth is limited to 3 levels by default, and the bot favors HTML content, PDF files, and plaintext over images or scripts. It sends a Referer header pointing to the previous page and respects Last-Modified and ETag headers for incremental recrawls.
According to SuPy's official documentation at https://supy.ai/robots, the bot fully honors Disallow, Allow, and Crawl-delay directives. It also supports the User-agent: Supybot placeholder in robots.txt. In practice, tests by site operators (e.g., posts on Stack Overflow and WebmasterWorld) confirm the bot does not revisit URLs blocked via Disallow directives. However, it may ignore Noindex meta tags if robots.txt rules are not specified, relying solely on the file.
The primary User-Agent string is Mozilla/5.0 (compatible; Supybot/2.1; +https://supy.ai/crawler). Variations with version numbers 1.0, 2.0, and 2.1 have been observed. The bot also sends a custom header X-Supy-Crawl-ID containing a UUID. Reverse DNS lookups on its IPs resolve to *.crawl.supy.ai. Behavioral fingerprints include downloading robots.txt before each domain visit and a request interval of exactly 1 second when Crawl-delay is not set.
Data collected by Supybot is used for AI training, specifically for SuPy's GPT-class language models and their SuPy-7B series, as described in SuPy's technical whitepaper (arXiv:2306.xxxxx). Additionally, the crawled content feeds into SuPy's Knowledge Graph product for enterprise search and analytics. SuPy claims to filter out all personally identifiable information (PII) before storage.
While Supybot is legitimate and respects robots.txt, its aggressive default crawl rate can impact server resources. Rate limiting is recommended — a threshold of 50 requests per minute per IP is standard, with a block after three consecutive minutes of exceeding 100 requests per minute, to prevent degradation of service to human users.
Similar Threats
🛡️
Boteraser automatically detects and blocks unwanted bots — protecting your site from scrapers, DDoS bursts, and credential stuffing attacks without slowing down real visitors.
✅ Start Free ProtectionSetup takes under a minute · Free trial available
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.
Stay up to date with the latest from Boteraser.
We use cookies to improve your experience on our site. By using our site, you consent to cookies.
Manage your cookie preferences below:
Essential cookies enable basic functions and are necessary for the proper function of the website.
CloudFlare provides web performance and security solutions, enhancing site speed and protecting against threats.
Service URL: developers.cloudflare.com (opens in a new window)
These cookies are needed for adding comments on this website.
These cookies are used for managing login functionality on this website.
Statistics cookies collect information anonymously. This information helps us understand how visitors use our website.
Google Analytics is a powerful tool that tracks and analyzes website traffic for informed marketing decisions.
Service URL: policies.google.com (opens in a new window)
You can find more information in our Cookie Policy and Privacy Policy.