GridBot

Bot User-Agent: gridbot

🤖 Overview

GridBot is a web crawler operated by Grid.ai, a platform that provides cloud infrastructure and tools for training machine learning models. First documented in their official crawler policy (grid.ai/bot-policy), the bot collects publicly available web content—including text, images, and metadata—to augment training datasets for Grid.ai’s foundational AI models, improving their performance across natural language and vision tasks.

🌐 Technical Behavior

GridBot performs both broad and targeted crawls, initiating requests from IP addresses within Amazon Web Services (AWS) EC2 ranges, as confirmed by Grid.ai’s documentation (docs.grid.ai/platform/crawler). The bot sends requests with a default delay of 1 second per domain but may increase frequency for high-quality sources. It uses HTTP/1.1 and respects 304 Not Modified responses to minimize bandwidth. The crawler identifies itself via the User-Agent header and provides a contact email in its From header. It follows links up to a configurable depth of 3–5 levels and indexes both HTML and linked resources such as PDFs or images. GridBot also sends Accept-Encoding: gzip headers to reduce transfer size.

📋 robots.txt Compliance

GridBot fully honors robots.txt directives, including Disallow and Crawl-Delay rules, as stated in Grid.ai’s bot policy page (grid.ai/bot-policy). The company explicitly notes that the crawler checks the file before each request and also respects noindex meta tags. Non-compliance with a site’s directives results in immediate termination of the crawl and removal from the bot’s queue.

🔍 Detection Indicators

The primary User-Agent string is GridBot/1.0 (optionally with (+https://grid.ai/bot) appended). Additional fingerprints include a consistent Accept header of text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8 and a From header set to [email protected]. The bot does not set a Referer header and uses static SSL certificates from AWS’s certificate manager, often from the us-east-1 region.

📊 Data Usage

Collected data is used exclusively for training Grid.ai’s proprietary machine learning models, including large language models and computer vision systems. Grid.ai states that no personal identifiable information (PII) is intentionally collected and that all data is anonymized before training (grid.ai/privacy). The data may also be shared with research partners under strict agreements to improve model robustness and fairness.

⚙️ Rate Limiting Policy

Webmasters should rate-limit GridBot to prevent excessive resource consumption, as the bot can become aggressive on popular domains, sending up to 5 requests per second during peak bursts. A threshold of 10 requests per minute per IP is recommended, with a 503 response or connection drop for violations. This policy balances the bot’s legitimate need for data with site stability and is consistent with industry best practices for non-malicious crawlers.

Free Traffic Analysis

What's Actually Crawling Your Website?

Discover which unwanted bots are being blocked on your site, how often they hit, and where they come from — real data from your own traffic, not guesswork.

🔍 Scan My Site Free

Powered by JA4 fingerprinting, honeypot traps & behavioral analysis

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.