SiteExplorer

Bot User-Agent: siteexplorer

🤖 Overview

SiteExplorer is a web crawler historically operated by Alexa Internet, a wholly owned subsidiary of Amazon that provided web analytics and competitive intelligence until its shutdown on May 1, 2022. The bot fed data into Alexa’s flagship Site Explorer tool, which offered backlink profiles, traffic estimates, and SEO metrics by crawling publicly accessible web pages. Alexa’s crawler infrastructure, including SiteExplorer, was used to build the company’s proprietary ranking system, the Alexa Traffic Rank, which evaluated website popularity based on combined visitor traffic from users of the Alexa Toolbar and other sources.

🌐 Technical Behavior

The SiteExplorer bot performed wide-scale web crawling at variable frequencies, typically requesting multiple pages per second from a given host during peak activity. It originated from IP addresses belonging to Amazon’s corporate ASN (AS16509), with ranges such as 54.x.x.x and 52.x.x.x, as documented in Alexa’s published IP lists. The crawler used HTTP/1.1 with standard GET requests and did not support persistent connections in its early versions. It respected the robots.txt Crawl-Delay directive if explicitly set, but often ignored noindex meta tags in practice, instead relying on further processing to exclude low-value pages. The bot’s crawl depth was moderate, usually limited to a few thousand URLs per site per crawl cycle, and it avoided crawling binary files such as images or PDFs unless linked from text content.

📋 robots.txt Compliance

Official documentation from Alexa confirmed that SiteExplorer honored robots.txt Disallow directives, provided they were not overly restrictive. A known exception was the bot’s disregard for the Crawl-Delay instruction when set below one second, as the crawler’s internal rate limiter could only throttle to a minimum of one request per second. This behavior was explicitly noted in Alexa’s developer support pages before the service shutdown.

🔍 Detection Indicators

The primary User-Agent string for SiteExplorer was “SiteExplorer/1.0”, sometimes preceded by “Mozilla/5.0 (compatible;)”. Additionally, the bot used the identifying header “From: [email protected] and the IP ranges were traceable via reverse DNS entries ending in .amazonaws.com. Behavioral fingerprints included a high request rate—often over 10 per second—combined with a lack of JavaScript execution and cookie support.

📊 Data Usage

The collected data powered Alexa’s Site Explorer dashboard, which displayed backlink graphs, referring domains, top keywords, and site ranking history. The crawled content was also used to feed Alexa’s search index and to compute the Alexa Traffic Rank, a metric based on toolbar data and panel estimates. After Amazon discontinued Alexa.com, the SiteExplorer crawler was decommissioned; the underlying dataset has been archived and is no longer updated.

⚙️ Rate Limiting Policy

SiteExplorer is rate-limited because its legacy crawling algorithm did not automatically back off under load, potentially overwhelming under-provisioned servers. A threshold-based block is justified to protect server resources while still allowing the bot to gather sufficient data for the now-defunct Alexa services.

🛡️

Stop Bots. Save Bandwidth. Protect Revenue.

Boteraser automatically detects and blocks unwanted bots — protecting your site from scrapers, DDoS bursts, and credential stuffing attacks without slowing down real visitors.

✅ Start Free Protection

Setup takes under a minute  ·  Free trial available

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.