SemrushBot-OCOB
Bot User-Agent:semrushbot-ocob
🤖 Overview
SemrushBot-OCOB is a legitimate web crawler operated by Semrush, a leading digital marketing platform headquartered in Boston, USA. Its primary purpose is to collect publicly accessible content from websites to build and update "Object" records—structured data about page elements, URLs, and content types—used in Semrush’s Site Audit, On-Page SEO Checker, and other SEO diagnostic tools. According to Semrush’s official bot documentation at https://www.semrush.com/bot/, this agent specifically focuses on extracting object-level data (e.g., headings, images, structured data) rather than full page text, supporting the company’s suite of competitive analysis and website optimization products.
🌐 Technical Behavior
SemrushBot-OCOB performs directed crawls starting from user-submitted URLs or discovered links, with a default request frequency of approximately 2–3 requests per second per domain, as documented in Semrush’s public bot policy. It uses HTTP/1.1 and respects Transfer-Encoding: chunked responses. The bot originates from a defined set of IPv4 ranges published at https://www.semrush.com/bot/ip-list.txt, which includes subnets such as 185.183.106.0/24 and 195.182.172.0/24. It also supports Accept-Encoding: gzip for efficient data transfer. The crawler does not execute JavaScript; it only inspects static HTML and embedded resources, making it detectable by its minimal request footprint.
📋 robots.txt Compliance
SemrushBot-OCOB fully obeys robots.txt directives, as stated in Semrush’s official guidelines. The bot reads the Disallow rules for the exact User-Agent token SemrushBot-OCOB and for the generic * wildcard, ensuring that blocked paths are never requested. Verified via the publicly available crawling policy at https://www.semrush.com/bot/, the bot also respects Crawl-delay directives when present, reducing its request rate to the specified minimum.
🔍 Detection Indicators
The primary User-Agent string is Mozilla/5.0 (compatible; SemrushBot-OCOB/1.2; +https://www.semrush.com/bot/). Additional behavioral fingerprints include the From HTTP header set to [email protected] and a User-Agent that always contains the phrase SemrushBot. No custom X-Robots-Tag headers are sent, but the bot reads X-Robots-Tag responses. These patterns are documented in Semrush’s own technical reference and confirmed by community analysis.
📊 Data Usage
Collected object-level data feeds directly into Semrush’s Site Audit tool, which detects broken links, duplicate content, missing meta tags, and structured data errors. It also populates the On-Page SEO Checker dashboard, allowing users to compare their page objects against top-ranking competitors. Data is retained for up to 90 days and is never used for AI model training or sold to third parties, per Semrush’s privacy policy (see https://www.semrush.com/company/legal/privacy/).
⚙️ Rate Limiting Policy
SemrushBot-OCOB is rate-limited by Semrush to ~2 requests per second per domain, but aggressive server-side throttling may be applied by site owners to prevent resource exhaustion. The policy rationale is that even legitimate crawlers should be constrained to avoid degrading service performance for real users, and threshold-based blocking is recommended when the bot exceeds documented crawl limits.
Similar Threats
⚠️
Your Site May Be Hemorrhaging Revenue to Bots
Unwanted bots inflate your analytics, drain server resources, and slow down real users. Check if your site is affected — completely free.
Check My Site for FreeFree to start · Cancel anytime
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.