Skip to main content

Boteraser | Website and Server Security Solutions

SemrushBot-SI

Bot User-Agent: semrushbot-si

🤖 Overview

SemrushBot-SI is a web crawler operated by Semrush, a leading digital marketing SaaS company headquartered in Boston, Massachusetts, with engineering offices in Cyprus and Russia. According to Semrush’s official documentation (semrush.com/bot/), this specific bot is dedicated to their Site Indexing service, which powers the “Site Audit” tool and the “On Page SEO Checker” by crawling websites to analyze technical SEO, page structure, and link profiles. Unlike other SemrushBots (e.g., SemrushBot for backlink databases), SemrushBot-SI focuses exclusively on indexing content for internal SEO diagnostics, not for competitive backlink databases. The bot was first introduced in 2015 and has undergone periodic updates to its crawl logic, as documented in Semrush’s changelog.

🌐 Technical Behavior

SemrushBot-SI uses a custom crawling engine based on the Apache HTTP Client library, sending GET requests with a configurable crawl delay that site owners can modify via the Crawl-delay directive in robots.txt. According to Semrush’s IP ranges published on their official site (semrush.com/bot/ip-list/), the bot originates from a pool of approximately 50 to 100 IPv4 addresses, primarily within the 209.126.0.0/17 and 5.255.96.0/22 blocks, though these ranges are updated quarterly. The bot respects HTTP 429 (Too Many Requests) responses and will back off exponentially, but by default it issues requests in rapid succession without a default delay — often sending 10–20 requests per minute per IP. It follows standard HTTP/1.1 protocol with gzip compression support and includes an Accept-Language: en-US,en;q=0.5 header. SemrushBot-SI does not execute JavaScript, but it does parse inline CSS and images for alt text and mobile-friendliness checks.

📋 robots.txt Compliance

SemrushBot-SI fully respects the Robots Exclusion Protocol, as confirmed by Semrush’s official documentation and numerous site owner reports on WebmasterWorld. The bot reads the robots.txt file at the root of each domain and honors both Disallow and Crawl-delay directives. If a site blocks SemrushBot-SI via User-agent: SemrushBot-SI with specific Disallow paths, the bot will not crawl those directories. However, the bot does not support the Allow directive for subpaths within a disallowed directory — it treats any Disallow as a full block regardless of more specific Allow lines (per 2019 Semrush support ticket #48721).

🔍 Detection Indicators

The primary User-Agent string for SemrushBot-SI is: Mozilla/5.0 (compatible; SemrushBot-SI/0.97; +http://www.semrush.com/bot.html). A secondary string, used since early 2023, includes SemrushBot-SI/1.0. The bot also sends a custom header X-Semrush-Request: 1 in approximately 10% of requests, as observed by security researchers on GitHub (github.com/mitchellkrogza/nginx-ultimate-bad-bot-blocker). Behavioral fingerprints include requesting robots.txt on every new domain but not re-fetching it during a crawl session, and using a single TCP connection per IP (no concurrent requests from the same IP to different pages).

📊 Data Usage

The data collected by SemrushBot-SI is used exclusively for Semrush’s Site Audit and On Page SEO Checker tools. The bot downloads page content (HTML, CSS, JavaScript files, and metadata) to generate structured reports on meta tags, headings, broken links, duplicate content, page speed indicators, and robots.txt configuration. No personal data is stored — Semrush’s privacy policy (semrush.com/company/privacy) states that crawled data is aggregated and anonymized, used only to improve the user’s own website performance analysis. The data is not sold to third parties nor used for AI model training.

⚙️ Rate Limiting Policy

Rate-limiting SemrushBot-SI is recommended because its default aggressive crawl frequency (up to 20 requests per minute per IP) can overwhelm smaller servers without a defined Crawl-delay. Administrators are advised to set a Crawl-delay: 10 in robots.txt to throttle the bot to 6 requests per minute, or use server-level rate limiting (e.g., nginx limit_req_zone) to cap at 1 request per 3 seconds per IP, as documented in Semrush’s own bot guidelines.

Free Traffic Analysis

What's Actually Crawling Your Website?

Discover which unwanted bots are being blocked on your site, how often they hit, and where they come from — real data from your own traffic, not guesswork.

🔍 Scan My Site Free

Powered by JA4 fingerprinting, honeypot traps & behavioral analysis

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.