Majestic12
Bot User-Agent:majestic12
🤖 Overview
Majestic12 (also known as the Majestic Bot) is a web crawler operated by Majestic, a UK-based SEO and digital marketing analytics company, now part of the Mx Group. Its primary purpose is to build and maintain the Majestic SEO backlink index, which maps the link graph of the public web, providing data on inbound links, anchor text, and trust flow metrics to subscribers. The bot was first documented in the early 2010s and has since evolved into one of the most aggressive yet legitimate crawlers for link analysis and site authority scoring.
🌐 Technical Behavior
The Majestic12 crawler operates with high concurrency, often sending multiple simultaneous requests to a single domain. It uses a custom HTTP client and supports both HTTP/1.1 and HTTP/2. Its crawl frequency can vary—on popular sites it may request hundreds of pages per hour—but it typically paces itself based on server response times. The bot resolves domain names and follows internal and external links recursively, building a directed graph representation of the web. IP ranges are drawn from Majestic’s own ASN (AS37235), with addresses in the 185.130.5.0/24 and 91.215.216.0/24 blocks, though these change over time. Official documentation at majestic.com/robots.txt and majestic.com/bot details its crawling strategy: it respects Crawl-Delay directives in robots.txt and avoids crawling pages with X-Robots-Tag: noindex or noarchive. It also supports If-Modified-Since headers to reduce redundant data transfers.
📋 robots.txt Compliance
According to Majestic’s own robotstxt guide (majestic.com/robotstxt), the bot fully honors Disallow directives found in a site’s robots.txt file. It also respects Crawl-Delay values; if no delay is specified, it defaults to a polite interval. There is no evidence of intentional violations—Majestic explicitly states that they block crawlers that ignore robots.txt. However, because the bot caches robots.txt results temporarily, changes to the file may take up to 24 hours to take effect.
🔍 Detection Indicators
The primary User-Agent string is Mozilla/5.0 (compatible; Majestic-12/1.0; +http://majestic12.co.uk/bot.php). Older versions used Majestic-12 without the Mozilla prefix. A secondary identifier is the X-Majestic-Referer header, which contains the source URL that led to the request. The bot also sends Accept: */* and Connection: keep-alive headers. Behavioral fingerprints include high request concurrency (often 10–50 simultaneous connections) and a strong preference for text/html content types.
📊 Data Usage
All collected data is aggregated into the Majestic Flow Metrics database, which includes Citation Flow and Trust Flow scores for billions of web pages. This data is sold as a subscription service to SEO professionals, marketers, and researchers. Majestic does not use the raw crawled content for AI training or model development—it exclusively extracts link structure, anchor text, and basic page metadata (title, description) to map the web’s interconnection topology.
⚙️ Rate Limiting Policy
Because the Majestic12 bot can deliver high request rates without external pacing, it is strongly recommended to apply rate limiting (e.g., 10–50 requests per minute per IP) to protect server resources. Majestic itself advises site owners to set a Crawl-Delay in robots.txt to control the bot’s speed; fair throttle thresholds align with standard polite crawler access, preserving site performance while still allowing valuable backlink indexing.
Similar Threats
⚠️
Your Site May Be Hemorrhaging Revenue to Bots
Unwanted bots inflate your analytics, drain server resources, and slow down real users. Check if your site is affected — completely free.
Check My Site for FreeFree to start · Cancel anytime
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.