Skip to main content

Boteraser | Website and Server Security Solutions

meta-externalagent

Bot User-Agent: meta-externalagent

🤖 Overview

meta-externalagent is a web crawler operated by Meta Platforms, Inc. (formerly Facebook), introduced around 2023 to collect publicly available web content for training large language models (LLMs) and improving Meta’s AI products such as LLaMA, Meta AI, and other generative AI services. According to Meta’s official developer documentation and their robots.txt file at facebook.com/robots.txt, the crawler is explicitly listed under the user-agent “meta-externalagent” with defined crawl rules.

🌐 Technical Behavior

The crawler sends HTTP GET requests with a default frequency of several requests per second, though Meta has not published exact rate limits. It originates from IP ranges registered to Meta (e.g., 31.13.24.0/21, 69.171.224.0/19, and others listed in Meta’s AS32934). The bot fetches both HTML and structured data (JSON-LD, Open Graph) and respects standard HTTP status codes (e.g., 429 for rate limiting). It uses TLS 1.2+ and follows redirects. Crawl patterns are breadth-first across domains, with a focus on high-quality, publicly accessible pages.

📋 robots.txt Compliance

Meta states in its official crawler documentation that meta-externalagent honors robots.txt Disallow directives. Evidence from Meta’s own robots.txt file shows it includes a specific entry for this agent under a dedicated user-agent block, allowing site owners to restrict access. No known cases of deliberate non-compliance have been reported by security researchers or webmasters.

🔍 Detection Indicators

The primary User-Agent string is: “Mozilla/5.0 (compatible; Meta-ExternalAgent/1.0; +https://developers.facebook.com/docs/sharing/bot/)”. Behavioral fingerprints include a consistent “From” header or “Referer” set to Meta’s domain, and a request pattern that includes an “Accept-Language: en-US” header. IP reverse DNS typically resolves to a *.fb.com or *.facebook.com hostname.

📊 Data Usage

Collected data is used exclusively for AI training and improving Meta’s language models, such as LLaMA, as well as for general knowledge base enrichment. Meta’s official documentation states that the crawler does not collect personal data without consent and follows privacy policies outlined at https://www.facebook.com/privacy/policy/.

⚙️ Rate Limiting Policy

This bot is rate‑limited by many webmasters because its crawl volume, while legitimate, can still consume significant server resources if left unchecked. A threshold-based rate limit (e.g., 10 requests per second per IP) is a reasonable security practice to prevent unintended load without blocking the bot entirely.

Free Traffic Analysis

What's Actually Crawling Your Website?

Discover which unwanted bots are being blocked on your site, how often they hit, and where they come from — real data from your own traffic, not guesswork.

🔍 Scan My Site Free

Powered by JA4 fingerprinting, honeypot traps & behavioral analysis

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.