meta-externalagent
meta-externalagent is a web crawler operated by Meta Platforms, Inc. (formerly Facebook), introduced around 2023 to collect publicly available web content for training large language models (LLMs) and improving Meta’s AI products such as LLaMA, Meta AI, and other generative AI services. According to Meta’s official developer documentation and their robots.txt file at facebook.com/robots.txt, the crawler is explicitly listed under the user-agent “meta-externalagent” with defined crawl rules.
The crawler sends HTTP GET requests with a default frequency of several requests per second, though Meta has not published exact rate limits. It originates from IP ranges registered to Meta (e.g., 31.13.24.0/21, 69.171.224.0/19, and others listed in Meta’s AS32934). The bot fetches both HTML and structured data (JSON-LD, Open Graph) and respects standard HTTP status codes (e.g., 429 for rate limiting). It uses TLS 1.2+ and follows redirects. Crawl patterns are breadth-first across domains, with a focus on high-quality, publicly accessible pages.
Meta states in its official crawler documentation that meta-externalagent honors robots.txt Disallow directives. Evidence from Meta’s own robots.txt file shows it includes a specific entry for this agent under a dedicated user-agent block, allowing site owners to restrict access. No known cases of deliberate non-compliance have been reported by security researchers or webmasters.
The primary User-Agent string is: “Mozilla/5.0 (compatible; Meta-ExternalAgent/1.0; +https://developers.facebook.com/docs/sharing/bot/)”. Behavioral fingerprints include a consistent “From” header or “Referer” set to Meta’s domain, and a request pattern that includes an “Accept-Language: en-US” header. IP reverse DNS typically resolves to a *.fb.com or *.facebook.com hostname.
Collected data is used exclusively for AI training and improving Meta’s language models, such as LLaMA, as well as for general knowledge base enrichment. Meta’s official documentation states that the crawler does not collect personal data without consent and follows privacy policies outlined at https://www.facebook.com/privacy/policy/.
This bot is rate‑limited by many webmasters because its crawl volume, while legitimate, can still consume significant server resources if left unchecked. A threshold-based rate limit (e.g., 10 requests per second per IP) is a reasonable security practice to prevent unintended load without blocking the bot entirely.
Similar Threats
Free Traffic Analysis
Discover which unwanted bots are being blocked on your site, how often they hit, and where they come from — real data from your own traffic, not guesswork.
🔍 Scan My Site FreePowered by JA4 fingerprinting, honeypot traps & behavioral analysis
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.
Stay up to date with the latest from Boteraser.
We use cookies to improve your experience on our site. By using our site, you consent to cookies.
Manage your cookie preferences below:
Essential cookies enable basic functions and are necessary for the proper function of the website.
CloudFlare provides web performance and security solutions, enhancing site speed and protecting against threats.
Service URL: developers.cloudflare.com (opens in a new window)
These cookies are needed for adding comments on this website.
These cookies are used for managing login functionality on this website.
Statistics cookies collect information anonymously. This information helps us understand how visitors use our website.
Google Analytics is a powerful tool that tracks and analyzes website traffic for informed marketing decisions.
Service URL: policies.google.com (opens in a new window)
You can find more information in our Cookie Policy and Privacy Policy.