GoogleOther-Image

Bot User-Agent: googleother-image

🤖 Overview

GoogleOther-Image is a web crawler operated by Google LLC, introduced as part of the GoogleOther family in early 2023. Unlike the primary Googlebot used for search indexing, GoogleOther-Image is dedicated to collecting publicly accessible image content specifically for Google’s internal machine learning projects, including training multimodal models such as those powering Google Gemini (formerly Bard) and other AI research initiatives. The crawler was first documented in Google’s official crawler documentation in February 2023, and its purpose is distinct from search index building—it gathers images to improve vision-language understanding, object recognition, and generative image capabilities.

🌐 Technical Behavior

GoogleOther-Image follows a consistent crawl pattern: it fetches image URLs referenced in HTML img tags and srcset attributes, as well as linked image files (e.g., JPEG, PNG, WebP). Requests originate from IP ranges within Google’s owned ASN (AS15169) and resolve to hostnames under .googlehosted.com and .google.com. Official documentation lists the typical request frequency as moderate, with a default crawl delay of approximately 1 request per second per host, though this may vary based on site responsiveness. The crawler uses HTTP/1.1 and HTTP/2 protocols, sends a User-Agent header of GoogleOther-Image, and includes Accept: image/webp,image/apng,image/*,*/*;q=0.8. It does not execute JavaScript and only fetches static image resources. Google’s guidelines note that the crawler may also request robots.txt before each domain visit to check for disallowed paths.

📋 robots.txt Compliance

According to Google’s official documentation (developers.google.com/search/docs/crawling-indexing/google-other-crawlers), GoogleOther-Image fully respects robots.txt directives when the user-agent line User-agent: GoogleOther-Image is used. If no specific rule is defined for that user-agent, it falls back to the wildcard rule (User-agent: *). However, Google notes that unlike Googlebot, this crawler is not used for search indexing and therefore does not respect the older Googlebot-Image directives—site owners must explicitly target GoogleOther-Image to block or allow image access. Evidence from public web server logs confirms that the crawler strictly obeys Disallow paths and does not attempt to bypass restrictions.

🔍 Detection Indicators

The primary detection fingerprint is the User-Agent string: Mozilla/5.0 (Linux; Android 6.0.1; Nexus 5X Build/MMB29P) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Mobile Safari/537.36 (compatible; GoogleOther-Image). This string is unique and consistent across all requests. Additionally, the crawler always includes a From: [email protected] header, though this is optional. Reverse DNS lookups on requesting IPs reveal hostnames matching *.googlebot.com or *.googleusercontent.com. Traffic typically originates from IP blocks like 66.249.64.0/19 and 216.58.192.0/19, as documented by Google’s published IP ranges (support.google.com/webmasters/answer/80553). Behavioral indicators include a consistent rate of one request per second, no support for If-Modified-Since headers, and a preference for HTTPS connections.

📊 Data Usage

The collected images are used exclusively for internal Google AI model training, not for search indexing or public display. Google states in its guidelines that GoogleOther-Image data feeds into machine learning datasets that improve product features like Google Lens, Google Photos smart categorization, and the image understanding capabilities of Gemini. The data is processed in Google’s data centers and retained per Google’s privacy policy—images may be used to train models that do not reproduce original works. No third-party access or sharing is documented.

⚙️ Rate Limiting Policy

Site owners are advised to rate-limit GoogleOther-Image using threshold-based blocking because, while legitimate, it can consume significant bandwidth if crawling large image libraries. A common policy is to set a request limit of 10 requests per second per IP, returning HTTP 429 (Too Many Requests) when exceeded. This prevents disruption to regular site operations while still allowing the crawler to operate within acceptable bounds.

🛡️

Stop Bots. Save Bandwidth. Protect Revenue.

Boteraser automatically detects and blocks unwanted bots — protecting your site from scrapers, DDoS bursts, and credential stuffing attacks without slowing down real visitors.

✅ Start Free Protection

Setup takes under a minute  ·  Free trial available

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.