meta-webindexer

Indexer User-Agent: meta-webindexer

🤖 Overview

meta-webindexer is a legitimate web crawler operated by Meta Platforms, Inc., first publicly documented in 2023, designed to index publicly accessible web content for integration into Meta’s search and recommendation systems across its platforms including Facebook, Instagram, and Threads. According to Meta’s official developer documentation (https://developers.facebook.com/docs/sharing/webmasters/crawler), the bot gathers data for improving product features such as search results and link previews, and may also contribute to training of Meta’s AI models like LLaMA under their data usage policy.

🌐 Technical Behavior

The meta-webindexer performs HTTP GET requests with standard crawling patterns, typically observing a crawl delay of 1‑2 seconds between requests as seen in server logs. Its requests originate from IP ranges registered to Meta, including 31.13.24.0/21 and 69.171.224.0/20, which are also used by other Meta crawlers like FacebookExternalHit. The bot uses HTTP/1.1 and HTTPS protocols, and includes a User-Agent header of Meta-WebIndexer/1.0 plus version variants. It fetches robots.txt before crawling a domain and adheres to Crawl-Delay directives. Publicly stated crawl rates approximate 100 requests per minute per domain, though this varies with site responsiveness.

📋 robots.txt Compliance

Meta’s official documentation explicitly states that meta-webindexer respects Disallow directives in robots.txt, as well as Allow and User-agent rules. However, the bot does not honor the nofollow meta tag or X-Robots-Tag header for indexing decisions; it relies solely on robots.txt for access control. This behavior is consistent with reports from webmasters in community forums.

🔍 Detection Indicators

Primary identification is via the User-Agent string Meta-WebIndexer/1.0 (with possible version variants). Additional headers may include From: [email protected] and a X-Purpose header sometimes set to preview or search. Log entries typically show reverse DNS records resolving to *.fb.com or *.meta.com.

📊 Data Usage

Collected data is primarily used to power Meta’s internal search engines across its social platforms, providing users with relevant web results and link summaries. Under Meta’s privacy policy, crawled content may also be used to train and improve large language models such as LLaMA and other AI systems, with Meta asserting that only public, non‑protected content is ingested.

⚙️ Rate Limiting Policy

While meta-webindexer is a legitimate bot, it can generate significant traffic if multiple pages are crawled simultaneously, and it is therefore rate‑limited by many webmasters. The recommended policy is to impose a threshold of 20 requests per 10 seconds per IP range, blocking further requests until the rate drops, in order to protect server resources without completely disabling indexing.

53% of Web Traffic Is Bots in 2026

— Imperva Bad Bot Report 2026

How much of your traffic is automated? Get your personal bot traffic report and see exactly what's hitting your server — completely free.

📊 Get My Bot Report

Sign up in seconds  ·  No card required

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.