geometabot
Geometabot is a web crawler operated by Meta Platforms, Inc. (formerly Facebook), used to collect publicly available web content for training and improving Meta’s large language models, including the LLaMA series. First publicly documented in early 2024 via Meta’s official crawler documentation, geometabot is one of several bots Meta employs to index the web for AI training datasets. The bot is explicitly distinguished from Meta’s other crawlers like FacebookExternalHit and mario-crawler, focusing specifically on AI model training rather than social graph or advertising signals.
Geometabot performs HTTP/1.1 and HTTP/2 GET requests over IPv4 and IPv6, primarily from IP ranges announced by Meta’s autonomous system AS32934 and AS54113. Observed crawl rates vary from a few requests per minute to several hundred per second during bulk operations, with a documented maximum crawl rate of approximately 100 requests per second per IP before throttling. The bot respects a configurable crawl delay via the Crawl-Delay directive in robots.txt, defaulting to one second if unspecified. Requests include a User-Agent header of geometabot/1.0 and optionally an Accept-Encoding: gzip header. Meta’s official guidance states geometabot does not execute JavaScript, nor does it follow meta refresh redirects, focusing solely on static HTML content. Its DNS PTR records often resolve to hostnames containing fwdproxy or any, reflecting Meta’s use of forward proxy infrastructure.
According to Meta’s official documentation published at https://developers.facebook.com/docs/sharing/bot/, geometabot fully honors Disallow directives in robots.txt. It also respects Allow, Crawl-Delay, and Sitemap directives. Meta explicitly advises site operators to block geometabot via robots.txt if they wish to exclude their content from AI training datasets, and the bot has been observed complying with such blocks in independent testing.
The primary identification string is the User-Agent token geometabot/1.0. Additional behavioral fingerprints include a consistent User-Agent header without accompanying vendor-specific comment strings (e.g., no +(http://...) annotation). The bot does not send a Referer header and uses a default Accept header of text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8. Reverse DNS lookups on source IPs from Meta’s AS32934 often reveal hostnames ending in .fbcdn.net or .tfbnw.net.
Data collected by geometabot is used exclusively for training Meta’s large language models, including LLaMA 2, LLaMA 3, and future variants. The content is processed to generate natural language training corpora, not for advertisement targeting or social media indexing. Meta states that private or sensitive personal data inadvertently collected is filtered through automated de‑identification pipelines before training use.
Rate limiting geometabot is recommended because the bot can issue thousands of requests per hour during deep crawls, potentially impacting server performance for shared or lower‑capacity hosting environments. A threshold of 10‑20 requests per second per IP is a common defensive measure that still allows the bot to gather data without overwhelming origin servers.
— Imperva Bad Bot Report 2026
How much of your traffic is automated? Get your personal bot traffic report and see exactly what's hitting your server — completely free.
📊 Get My Bot ReportSign up in seconds · No card required
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.
Stay up to date with the latest from Boteraser.
We use cookies to improve your experience on our site. By using our site, you consent to cookies.
Manage your cookie preferences below:
Essential cookies enable basic functions and are necessary for the proper function of the website.
CloudFlare provides web performance and security solutions, enhancing site speed and protecting against threats.
Service URL: developers.cloudflare.com (opens in a new window)
These cookies are needed for adding comments on this website.
These cookies are used for managing login functionality on this website.
Statistics cookies collect information anonymously. This information helps us understand how visitors use our website.
Google Analytics is a powerful tool that tracks and analyzes website traffic for informed marketing decisions.
Service URL: policies.google.com (opens in a new window)
You can find more information in our Cookie Policy and Privacy Policy.