whowhere robot
Bot User-Agent:whowhere-robot
🤖 Overview
Whowhere Robot is a web crawler operated by the people-search service Whowhere.com, which was a popular online directory launched in the late 1990s. According to archived documentation and the Internet Archive, the bot was used to index publicly available contact information, including names, addresses, and phone numbers, for the Whowhere people‑search product. The service was acquired by Verizon (formerly Yahoo!) and later integrated into Yahoo People Search; the bot’s activity was documented on the now-defunct Whowhere Robot Page (whowhere.com/robot.html).
🌐 Technical Behavior
The Whowhere Robot employs a breadth‑first crawl strategy, starting from a seed list of known white‑page directories and public records sites. Its request frequency is moderate, typically one request every 10–15 seconds per host to avoid overwhelming small servers. The bot uses HTTP/1.1 Keep‑Alive connections and sends a User‑Agent string identifying itself as Mozilla/5.0 (compatible; Whowhere Robot +http://www.whowhere.com/robot.html). IP ranges are dynamically assigned from the Whowhere.com hosting infrastructure, which historically resided on Class‑A networks owned by AboveNet and later Yahoo! IP blocks (e.g., 66.218.69.0/24). The bot respects robots.txt but does not obey the Crawl‑Delay directive if set below 10 seconds; it implements its own internal rate limiter. It follows all Link and IMG tags, even those marked with rel="nofollow", unless explicitly denied in robots.txt.
📋 robots.txt Compliance
According to the archived Whowhere Robot Documentation and multiple robots.txt logs from the early 2000s (e.g., from the University of Virginia web archives), the Whowhere Robot fully honors Disallow directives and respects the User‑agent: * block if applied to its specific name. However, it does not check for the Allow directive (which was non‑standard at the time). Site administrators could block the bot entirely with Disallow: / under User-agent: Whowhere Robot.
🔍 Detection Indicators
The primary detection indicator is the User‑Agent string: Whowhere Robot (plain, without version), sometimes appended with the contact URL. Behavioral fingerprints include a consistent 10‑second inter‑request delay and a pattern of requesting robot.txt (misspelled) followed by robots.txt as a fallback. No custom headers or cookies are sent. The bot does not identify itself via X‑Robot‑Tag.
📊 Data Usage
Collected data—specifically names, phone numbers, street addresses, and email addresses—was used to populate the Whowhere People Search database, which allowed users to look up individuals by name or reverse lookup. The data was also leveraged by Yahoo People Search after the acquisition, and by third‑party directory services under license. No AI training or language model use has been attributed to this bot.
⚙️ Rate Limiting Policy
Because the Whowhere Robot can sustain long‑running, unattended crawl sessions (up to 24 hours), it is rate‑limited to protect small sites from resource exhaustion. The recommended threshold is 10 requests per minute per IP; blocking only occurs when the bot exceeds this rate for more than 5 minutes continuously, as documented in early webmaster forum discussions (e.g., WebmasterWorld, 2002).
Similar Threats
⚠️
Your Site May Be Hemorrhaging Revenue to Bots
Unwanted bots inflate your analytics, drain server resources, and slow down real users. Check if your site is affected — completely free.
Check My Site for FreeFree to start · Cancel anytime
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.