myfamilybot

Bot User-Agent: myfamilybot

🤖 Overview

myfamilybot is a legitimate web crawler operated by MyFamily.com (now part of AncestryDNA LLC), a family‑history and social‑networking platform launched in the early 2000s. Its primary purpose is to index publicly available web content—such as family‑tree pages, obituaries, and genealogical records—to feed into MyFamily’s search and recommendation engine, enabling users to discover familial connections and historical records. Official documentation from Ancestry (ancestry.com/robots.txt) confirms its presence as a standard crawler.

🌐 Technical Behavior

The bot follows a conservative crawl pattern, typically requesting pages at intervals of 10–20 seconds to avoid overwhelming servers. It uses HTTP/1.1 and respects Last‑Modified headers to avoid re‑fetching unchanged content. IP ranges are assigned from Ancestry’s public ASN (AS 16509, Amazon Web Services’ infrastructure); crawls originate from IPv4 addresses in the 52.0.0.0/8 and 34.0.0.0/8 blocks. The bot primarily fetches HTML pages and ignores images, CSS, and JavaScript unless explicitly required for data extraction. It does not execute client‑side scripts. Crawl frequency varies by site popularity; high‑traffic domains may see multiple requests per day, but myfamilybot slows down upon receiving non‑200 responses.

📋 robots.txt Compliance

Based on publicly available robots.txt files (e.g., from Reuters, Wikipedia), myfamilybot fully honors Disallow directives. The official user‑agent string is “myfamilybot” (case‑insensitive), and it only crawls paths not explicitly blocked. There are no documented instances of deliberate violations; the bot respects crawl delays and Crawl‑Delay directives when specified.

🔍 Detection Indicators

The standard User‑Agent string is Mozilla/5.0 (compatible; myfamilybot/1.0; +https://www.myfamily.com/help/crawler). It may also appear as “myfamilybot” without the Mozilla prefix. Behavioral fingerprints include sequential GET requests, a low request rate, and a referrer header of “https://www.myfamily.com/”. No additional proprietary headers are known.

📊 Data Usage

Collected data—spanning public family‑tree pages, obituary texts, and genealogical forum posts—is indexed to power MyFamily’s internal search and automated family‑connection suggestions. The platform does not train commercial AI models on this data; it is used exclusively for personal family‑history discovery. Data is retained for the duration of active indexing and is deleted upon user request under Ancestry’s privacy policy.

⚙️ Rate Limiting Policy

This bot is rate‑limited because, while legitimate, it can generate non‑trivial traffic on low‑bandwidth servers. A threshold of 50 requests per minute per IP is recommended; blocking is never warranted, but throttling ensures fair resource allocation for all crawlers.

Free Traffic Analysis

What's Actually Crawling Your Website?

Discover which unwanted bots are being blocked on your site, how often they hit, and where they come from — real data from your own traffic, not guesswork.

🔍 Scan My Site Free

Powered by JA4 fingerprinting, honeypot traps & behavioral analysis

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.