grapefx

Bot User-Agent: grapefx

🤖 Overview

GrapeFX is a web crawler operated by Grape Data Inc., first publicly documented in a 2021 blog post at https://grape.io/blog/introducing-grapefx, designed to collect publicly accessible web content for training proprietary large language models and improving natural language processing systems. The crawler feeds data into Grape's AI product suite, including their text generation platform and enterprise analytics tools, as described in their official technology overview at https://grape.io/technology.

🌐 Technical Behavior

GrapeFX employs a distributed crawling architecture using IP addresses from major cloud providers such as AWS (EC2 regions us-east‑1, eu‑west‑1) and Google Cloud (us‑central1, europe‑west2), with requests originating from a broad range of /24 subnets listed in the Grape Data ASN (AS394512) as verified by BGP‑tool data. The bot averages 10‑20 requests per second per IP, with bursts up to 50 rps during initial discovery phases, and supports HTTP/1.1 with keep‑alive connections and HTTP/2 multiplexing. It respects the Accept‑Encoding header for gzip and brotli compression, uses a rotating User‑Agent string that appends a unique identifier per crawl session, and follows all links as well as sitemaps as per the sitemaps.org protocol, as detailed in their official developer documentation at https://grape.io/developers/crawler-specs.

📋 robots.txt Compliance

Per Grape Data’s public crawling policy at https://grape.io/robots-policy, GrapeFX honors both Disallow and Crawl‑Delay directives in robots.txt and reduces its request rate accordingly. However, a 2023 security advisory (GRA‑2023‑001) noted a brief period where the bot incorrectly interpreted wildcard patterns; this was patched in version 1.2.4 and is no longer active. The bot also respects X‑Robots‑Tag HTTP headers and meta tags as outlined in their changelog at https://github.com/grape-data/crawler/releases/tag/v1.2.4.

🔍 Detection Indicators

The primary User‑Agent string is Mozilla/5.0 (compatible; GrapeFX/1.0; +https://grape.io/bot), occasionally appearing as GrapeFX/1.0 (bot; [email protected]). Behavioral fingerprints include a missing Referer header on over 95% of requests, a consistent header order (User‑Agent, Accept, Accept‑Encoding, Host), and the presence of a custom X‑Grape‑Crawl header set to "1" in version 1.3+ (documented at https://grape.io/headers). IP ranges can be found in public IP‑lists maintained by Grape Data on their GitHub repository at https://github.com/grape-data/ip-lists.

📊 Data Usage

Collected data is used primarily for AI training to improve Grape’s language models, including the Grape‑LLM series, as well as for building structured knowledge graphs and analytics dashboards. According to their privacy policy at https://grape.io/privacy, all crawled data is anonymized, stripped of personal identifiable information, and not redistributed publicly. The data also feeds into Grape’s internal search indexing for their product documentation portal.

⚙️ Rate Limiting Policy

While GrapeFX is a legitimate crawler, its aggressive crawling patterns—especially during initial discovery—can cause server load spikes, making rate limiting necessary. Webmasters should implement throttling thresholds (e.g., block if requests exceed 30 per minute from a single IP) to ensure performance, while still allowing the bot reasonable access for data collection, consistent with recommendations from Grape’s own rate‑limiting guidelines at https://grape.io/rate-limits.

🛡️

Stop Bots. Save Bandwidth. Protect Revenue.

Boteraser automatically detects and blocks unwanted bots — protecting your site from scrapers, DDoS bursts, and credential stuffing attacks without slowing down real visitors.

✅ Start Free Protection

Setup takes under a minute  ·  Free trial available

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.