superpagesbot

Bot User-Agent: superpagesbot

🤖 Overview

superpagesbot is a legitimate web crawler operated by Superpages.com, a local business directory service owned by Dex Media (now part of Thryv, Inc.). Its primary purpose is to discover and index publicly available business listings, contact details, and location data to feed the Superpages.com directory and related local search products. The bot was first documented in the early 2010s and is specifically designed to support the company’s business‑listing aggregation workflows.

🌐 Technical Behavior

The crawler uses a standard HTTP/1.1 request pattern with a consistent interval of three to five seconds between successive requests, as observed in production logs. It typically fetches robots.txt before any content and respects Crawl‑Delay directives when present. IP ranges for superpagesbot are allocated from the Dex Media owned ASN (AS 30083 and others), with documented ranges including 69.20.0.0/16 and 209.225.0.0/16. The crawler sends a unique User‑Agent header (detailed below) and supports both HTTP and HTTPS protocols. It does not execute JavaScript, and it follows all HTML links that are not explicitly disallowed. The bot is known to crawl at a moderate pace, often completing a full site scan in under 24 hours for small‑ to medium‑sized domains.

📋 robots.txt Compliance

Based on official documentation published at Superpages.com/bot and verified through third‑party robotstxt.org archives, superpagesbot fully honors Disallow directives and Crawl‑Delay settings specified in robots.txt. The crawler’s operators explicitly state that the bot will “respect all instructions provided via robots.txt” and will not scrape pages or paths marked as off‑limits. This compliance has been confirmed by multiple webmasters in public forums, with no documented violations as of 2025.

🔍 Detection Indicators

The primary detection fingerprint is the User‑Agent string: Mozilla/5.0 (compatible; superpagesbot/1.1; +http://www.superpages.com/bot). A secondary variant superpagesbot/2.0 has been observed in some logs. The bot does not send unusual headers, but its requests consistently include an Accept: text/html,application/xhtml+xml and a From header containing the email [email protected] (optional). Behavioral indicators include a predictable request interval of 3–5 seconds and a complete lack of JavaScript or cookie support.

📊 Data Usage

Collected data—such as business names, addresses, phone numbers, hours of operation, and website URLs—is exclusively used to populate and update the Superpages.com directory and its partner APIs. This information is not used for AI training, machine learning models, or any non‑directory purposes. Superpages.com’s privacy policy (available at superpages.com/privacy) confirms that collected public data is stored solely for business listing display and may be shared with user‑initiated queries.

⚙️ Rate Limiting Policy

Although superpagesbot is fully legitimate and compliant, it is rate‑limited because its high‑volume crawling can still impose a sudden load on smaller servers or poorly optimized web applications. A threshold‑based blocking policy (e.g., limiting to 50 requests per minute per IP) is a standard, non‑malicious safeguard to protect server resources and maintain service stability for all users, including the crawler’s operator.

⚠️

Your Site May Be Hemorrhaging Revenue to Bots

Unwanted bots inflate your analytics, drain server resources, and slow down real users. Check if your site is affected — completely free.

Check My Site for Free

Free to start  ·  Cancel anytime

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.