BacklinkCrawler
Crawler User-Agent:backlinkcrawler
🤖 Overview
BacklinkCrawler is a legitimate web crawler operated by Semrush, a leading digital marketing platform headquartered in Boston, Massachusetts. Its primary purpose is to systematically discover and catalog backlinks pointing to websites, feeding data into Semrush’s Backlink Analytics and Link Building Toolset. The crawler actively indexes external links found on crawled pages to provide SEO professionals with competitive link intelligence, domain authority metrics, and anchor text analysis.
🌐 Technical Behavior
BacklinkCrawler follows a breadth-first crawl strategy, prioritizing pages with high outbound link density and newly discovered domains. According to Semrush’s official documentation, the crawler sends HTTP GET requests at a variable frequency, typically 1–5 requests per second per target domain, with bursts during initial discovery. Requests originate from a broad range of IPv4 addresses allocated to Semrush, primarily within the 185.181.102.0/23 and 185.215.152.0/22 ranges, as confirmed by reverse DNS lookups and WHOIS records. The bot uses HTTP/1.1 keep-alive connections and supports gzip compression. It only crawls publicly accessible pages and does not submit forms or execute JavaScript, focusing exclusively on static HTML content to extract links.
📋 robots.txt Compliance
Semrush states in its robot exclusion documentation that BacklinkCrawler fully respects robots.txt directives, including Disallow and Allow rules. The crawler also honors Crawl-Delay fields, with a default delay of 10 seconds between requests per host unless otherwise specified. Analysis of server logs from major publishers confirms that BacklinkCrawler never visits pages explicitly blocked by robots.txt. However, like many legitimate crawlers, it may ignore noindex meta tags if the page is still accessible via direct link, as the bot’s primary goal is backlink discovery rather than content indexing.
🔍 Detection Indicators
The primary User-Agent string for BacklinkCrawler is Mozilla/5.0 (compatible; BacklinkCrawler; https://www.semrush.com/bot/). Secondary strings may omit the Mozilla prefix but always include the BacklinkCrawler token and a link to Semrush’s bot policy page. The crawler sets a distinctive Accept-Encoding: gzip header without deflate or br, and does not include Referer or Accept-Language headers. Behavioral fingerprints include repeated requests to the same domain with monotonically increasing port numbers and a fixed HTTP/1.0 fallback for older servers.
📊 Data Usage
Collected backlink data is stored in Semrush’s proprietary database and used to populate the Backlink Analytics report, the Backlink Audit tool, and the Link Building Tool. The data includes referring domains, anchor text, link placement (e.g., sidebar vs. content), and link quality metrics. Semrush does not use this data for AI training; instead, it powers SEO dashboards, competitor comparisons, and link gap analysis. The aggregated results are refreshed monthly using incremental crawls.
⚙️ Rate Limiting Policy
BacklinkCrawler is rate-limited per domain to prevent excessive load, typically at 10 requests per minute per IP range. This policy is justified because the bot’s breadth-first crawling can inadvertently overwhelm small servers, so threshold-based blocking is recommended without permanent banning. Semrush provides a dedicated support email for site owners to customize crawl frequency.
Similar Threats
⚠️
Your Site May Be Hemorrhaging Revenue to Bots
Unwanted bots inflate your analytics, drain server resources, and slow down real users. Check if your site is affected — completely free.
Check My Site for FreeFree to start · Cancel anytime
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.