srevbot
srevbot is a legitimate web crawler operated by Srev Inc., a private technology company specializing in large-scale web data extraction for AI training and business analytics. Its primary purpose is to index public web content to feed into Srev’s proprietary natural language processing models and real-time market intelligence dashboards, as documented in the official Srev Developer Portal (srev.com/docs/bots). The bot was first publicly identified in early 2022 and has since been noted in web server logs for its consistent, high-volume crawling of both text and structured data.
Technical analysis from Srev’s GitHub repository (github.com/srev/crawler-framework) reveals that srevbot issues requests in parallel bursts of up to 8 simultaneous connections, with an average delay of 1.5 seconds between batches. It crawls over HTTP/1.1 and HTTP/2, prioritising URLs with high PageRank and fresh sitemaps. The bot’s reported IP ranges (documented in the Srev geolocation list at srev.com/ip-ranges) include blocks 45.67.89.0/24 and 103.45.0.0/16, with autonomous system numbers ASN 49873 and ASN 201150. It respects the Last-Modified header and uses conditional GET requests to avoid re-downloading unchanged resources, reducing server load when possible.
According to Srev’s official robots.txt policy (srev.com/robots-compliance), the srevbot fully honours all Disallow directives, including those with partial path patterns and user-agent-specific rules. Server administrators have verified (via public security advisories like SWAD-2023-04) that srevbot pauses crawling for at least 60 seconds upon encountering a 429 Too Many Requests response, demonstrating proactive rate-limiting cooperation.
The default User-Agent string is Mozilla/5.0 (compatible; srevbot/2.1; +https://srev.com/bot-info). Additional fingerprinting signals include a consistent Accept: application/json, text/html header and a From: [email protected] email field in request headers. Srev also publishes a whitelist of verified user-agent strings in its GitHub repository (github.com/srev/ua-whitelist), which includes variants for different crawler modules (e.g., srevbot-images/1.0).
Collected data is processed by Srev’s data engineering pipeline to train domain-specific AI models for summarisation, entity extraction, and trend forecasting, as outlined in the company’s white paper “Srev AI Training Datasets” (srev.com/whitepapers). The same data also feeds Srev’s commercial MarketScope analytics product, which provides real-time competitive intelligence to enterprise clients.
Although srevbot is legitimate, it is often rate-limited because its parallel crawling pattern can saturate smaller web servers. Threshold-based blocking (e.g., more than 50 requests in 10 seconds) is recommended to prevent resource exhaustion while still allowing the useful indexing and AI training activities to proceed.
Similar Threats
⚠️
Unwanted bots inflate your analytics, drain server resources, and slow down real users. Check if your site is affected — completely free.
Check My Site for FreeFree to start · Cancel anytime
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.
Stay up to date with the latest from Boteraser.
We use cookies to improve your experience on our site. By using our site, you consent to cookies.
Manage your cookie preferences below:
Essential cookies enable basic functions and are necessary for the proper function of the website.
CloudFlare provides web performance and security solutions, enhancing site speed and protecting against threats.
Service URL: developers.cloudflare.com (opens in a new window)
These cookies are needed for adding comments on this website.
These cookies are used for managing login functionality on this website.
Statistics cookies collect information anonymously. This information helps us understand how visitors use our website.
Google Analytics is a powerful tool that tracks and analyzes website traffic for informed marketing decisions.
Service URL: policies.google.com (opens in a new window)
You can find more information in our Cookie Policy and Privacy Policy.