GrapeshotCrawler
Crawler User-Agent:grapeshotcrawler
🤖 Overview
GrapeshotCrawler is a web crawler operated by Oracle Data Cloud (formerly Grapeshot Ltd., a UK-based company acquired by Oracle in 2019). Its primary purpose is to analyze publicly accessible web page content for contextual intelligence, enabling targeted advertising, brand safety verification, and audience segmentation. The crawler feeds data into Oracle's Contextual Intelligence and Data Management Platform (DMP) products, which categorize web pages by topics, sentiments, and keywords in real time.
🌐 Technical Behavior
The crawler operates with a moderate to aggressive crawl rate, often sending multiple requests per second from a distributed set of IP addresses. According to Oracle's official documentation, GrapeshotCrawler respects HTTP 429 (Too Many Requests) responses and will back off if rate-limited. Its IP ranges are not publicly listed in a single range but are known to originate from Oracle Cloud infrastructure and third-party data centers across the US, Europe, and Asia. The crawler uses HTTP/1.1 and HTTPS protocols and typically fetches only HTML pages and linked resources (CSS, JS) to evaluate content, but does not download images or videos. It also supports gzip compression for efficiency.
📋 robots.txt Compliance
Oracle explicitly states that GrapeshotCrawler honors Disallow directives in robots.txt and also respects X-Robots-Tag HTTP headers. However, there have been community reports of the crawler occasionally ignoring cached or stale robots.txt files; Oracle recommends webmasters to use robots.txt with a cache-control of 3600 seconds. Verified through Oracle’s official support pages and third-party audits.
🔍 Detection Indicators
The primary User-Agent string is Mozilla/5.0 (compatible; GrapeshotCrawler/2.0; +http://www.grapeshot.co.uk/) or GrapeshotCrawler/2.0. A secondary variant uses GrapeshotCrawler/3.0. Behavioral fingerprints include a high request frequency to a single domain within seconds, absence of JavaScript rendering, and a Via header sometimes containing “Oracle-Data-Cloud”. The crawler does not set a Referer header and typically begins crawling from seed URLs listed in its internal database.
📊 Data Usage
Collected data—primarily page titles, meta descriptions, headings, body text, and keyword density—is processed into contextual taxonomies used for programmatic ad placement. Oracle uses this information to match ads to relevant content without tracking individual user behavior, thereby supporting privacy-compliant advertising. The data is also aggregated into anonymized trend reports for publishers and advertisers.
⚙️ Rate Limiting Policy
Rate limiting is recommended because GrapeshotCrawler can generate significant server load—especially during initial crawls of large sites. Many webmasters throttle it to under 10 requests per second to prevent resource exhaustion. Oracle itself advises that the crawler will respect 429 responses, making threshold-based blocking the standard approach.
Similar Threats
53% of Web Traffic Is Bots in 2026
— Imperva Bad Bot Report 2026
How much of your traffic is automated? Get your personal bot traffic report and see exactly what's hitting your server — completely free.
📊 Get My Bot ReportSign up in seconds · No card required
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.