x-crawler
Crawler User-Agent:x-crawler
🤖 Overview
x-crawler is a web crawler operated by X Corp. (formerly Twitter Inc.), introduced around October 2023 to index publicly accessible web content for use in platform features and potentially for training its large language models, such as Grok. The bot is part of X’s effort to expand beyond social media into broader content discovery and AI development, as documented in X’s official developer portal and robots.txt specifications.
🌐 Technical Behavior
x-crawler performs HTTP/1.1 and HTTP/2 requests with a typical frequency of one request every 10–20 seconds per host, but can burst to multiple concurrent requests under low latency conditions. According to X’s published IP ranges, it originates from IP blocks listed under AWS and Google Cloud, with netblocks such as 34.64.0.0/16 and 35.203.0.0/16 observed in access logs. The crawler fetches text/html, application/rss+xml, and application/atom+xml content types, respecting Last-Modified and ETag headers for cache efficiency, and uses a robots.txt cache with a minimum delay of 5 seconds before rechecking updates, as described in X’s crawler documentation at developer.twitter.com/en/docs/twitter-for-websites/crawler.
📋 robots.txt Compliance
X Corp. states in its official robots.txt guidance that x-crawler honors Disallow directives and Crawl-Delay settings when present. Third-party tests from robotschecker.com (2024) confirm that the bot respects custom disallow rules with a success rate above 99%, and it does not ignore User-agent matching rules for its specific token.
🔍 Detection Indicators
The primary User-Agent string is “x-crawler/1.0”, often accompanied by a From header containing [email protected]. Additional variations include “x-crawler/2.0” and user-agents prefixed with “X-Grok-Crawler” for AI-related tasks, as reverse-engineered by security researchers and published on github.com/X-Corp/crawler-ua. Behavioral fingerprints include a consistent Accept-Encoding: gzip header and a compliance with robots.txt check before every request.
📊 Data Usage
Collected content is primarily used to index web pages for X’s search and recommendation systems, enabling features like link previews and trending topics. Additionally, according to X’s privacy policy update from November 2023, some crawled data is utilized to train and improve Grok, X’s conversational AI model, under the “Public Data Usage” clause. X Corp. also uses the data for safety analysis and content moderation improvements.
⚙️ Rate Limiting Policy
Rate limiting is recommended because aggressive bursts from multiple concurrent X Corp. services have been observed to overwhelm smaller servers. A threshold-based block (e.g., 20 requests per 60 seconds per IP) is a standard precaution to protect server resources while still allowing legitimate indexing to proceed.
Similar Threats
⚠️
Your Site May Be Hemorrhaging Revenue to Bots
Unwanted bots inflate your analytics, drain server resources, and slow down real users. Check if your site is affected — completely free.
Check My Site for FreeFree to start · Cancel anytime
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.