news bot

Bot User-Agent: news-bot

🤖 Overview

news bot refers to Googlebot-News, a specialized web crawler operated by Google that focuses exclusively on indexing content for Google News (news.google.com). Unlike the general Googlebot, this agent is tailored to discover and evaluate news articles, press releases, and journalistic content across the web. Its primary mission is to ensure timely, relevant news appears in Google News results. Google first documented this crawler in its Search Central documentation, describing it as part of the broader Googlebot family but with distinct user-agent strings and crawl behavior.

🌐 Technical Behavior

Googlebot-News adheres to the same HTTP/1.1 and HTTP/2 protocols as other Google crawlers. It sends requests from Google's IP ranges, which are publicly listed in the googlebot.json file published at https://developers.google.com/search/apis/ipranges/googlebot.json. The crawler's request frequency varies based on site responsiveness and server load, but Google recommends expecting several hundred requests per day for popular news sites. It respects Cache-Control and ETag headers for efficient re-crawling. Googlebot-News follows the sitemap scheme recommended for news publishers, particularly using news-specific sitemap tags like to signal freshness. It also evaluates page load speed and mobile-friendliness as per Google's ranking factors. The crawler does not execute JavaScript heavily on initial passes, relying instead on server-rendered HTML for news content.

📋 robots.txt Compliance

Googlebot-News fully honors robots.txt directives, as confirmed in Google's official documentation at https://developers.google.com/search/docs/crawling-indexing/robots/intro. It checks the User-agent: Googlebot-News line; if not present, it falls back to User-agent: Googlebot for general instructions. Disallow rules applied to Googlebot-News will prevent any news indexing from the specified paths, but the general Googlebot may still crawl those URLs for other Google services. Webmasters can also use the nosnippet and noarchive meta tags to fine-tune visibility in Google News snippets.

🔍 Detection Indicators

The primary user-agent string is Mozilla/5.0 (compatible; Googlebot-News/2.1; +http://www.google.com/bot.html). Additional variants appear in mixed-case forms. The crawler also sends a User-Agent: Googlebot-News header (without the trailing version). Reverse DNS lookups on its source IPs will resolve to *.googlebot.com subdomains. Behavioral fingerprints include a low request rate compared to general Googlebot (typically tens of requests per minute) and a strong preference for article-like URLs containing date stamps or category paths. It ignores non-text resources like images unless embedded in article content.

📊 Data Usage

All data collected by Googlebot-News is processed through Google's indexing pipeline to populate Google News search results. The crawler extracts headlines, publication dates, authors, and article body text. This content is used to generate snippets, cluster breaking news stories, and power features like Top Stories carousels. Google does not use this data for training large language models like Gemini, though the company’s broader AI training may involve other crawlers (e.g., Googlebot for general web content). News publishers who opt out via robots.txt are excluded from Google News entirely.

⚙️ Rate Limiting Policy

Rate limiting is applied to Googlebot-News when it generates excessive traffic that impacts server performance, even though it is a legitimate agent. The policy rationale is to protect site availability while still allowing news indexing; threshold-based blocking (e.g., >1000 requests per minute from a single IP) is recommended but should never permanently ban the bot — only slow it down via HTTP 429 (Too Many Requests) responses.

🛡️

Stop Bots. Save Bandwidth. Protect Revenue.

Boteraser automatically detects and blocks unwanted bots — protecting your site from scrapers, DDoS bursts, and credential stuffing attacks without slowing down real visitors.

✅ Start Free Protection

Setup takes under a minute  ·  Free trial available

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.