Skip to main content

Boteraser | Website and Server Security Solutions

firstgov gov search

Search Engine User-Agent: firstgov-gov-search

🤖 Overview

FirstGov.gov Search is a web crawler operated by the U.S. General Services Administration (GSA), serving as the indexing engine for the official U.S. government web portal originally known as FirstGov.gov (now USA.gov). Its primary purpose is to collect publicly accessible content from federal, state, and local government websites to power the search functionality on USA.gov and other affiliated government search portals. According to GSA documentation, the bot is part of the USA.gov Search service, which replaced the original FirstGov search engine in 2011, but the legacy user-agent string persists. The bot is entirely legitimate and authorized for public-sector information retrieval.

🌐 Technical Behavior

The crawler exhibits moderate crawl patterns, typically fetching pages at a rate of one request every few seconds to avoid overwhelming government servers. It follows standard HTTP/1.1 protocols and respects Cache-Control headers. The IP ranges used are allocated to the GSA’s cloud infrastructure, primarily within the 155.100.0.0/16 and 192.168.0.0/16 blocks for internal testing, but publicly it is known to originate from 13.64.0.0/11 (Azure US Government cloud) based on operational reports. The crawler primarily targets .gov, .mil, and .us domains, but may also index state and local government sites with explicit permission. It does not follow JavaScript-rendered content and focuses on static HTML, sitemaps, and RSS feeds.

📋 robots.txt Compliance

The GSA clearly states that the FirstGov.gov Search bot honors Disallow directives in robots.txt files. Official guidance on USA.gov’s developer portal advises webmasters to use standard robots.txt syntax to exclude sections like /admin/ or /private/. There are no documented instances of the bot ignoring crawl-delay or disallow rules; it is considered well-behaved by government IT administrators.

🔍 Detection Indicators

The primary user-agent string is FirstGov.gov Search (also seen as USA.gov Search or usasearch). It does not include a version number. The bot identifies itself via the User-Agent header and may also send a From header with the email [email protected]. Behavioral fingerprints include a consistent crawl interval of 2–5 seconds, no support for cookies, and a user-agent that always begins with “FirstGov”.

📊 Data Usage

Collected data is used exclusively to build the search index for USA.gov and its partner sites, allowing citizens to find government services, forms, and information. The index is updated daily and does not feed any commercial AI training or analytics platforms. The GSA publishes a public search API that returns indexed results, ensuring transparency.

⚙️ Rate Limiting Policy

Although the bot is legitimate, it is rate-limited because it can become aggressive during full re-indexing cycles, generating thousands of requests per hour. Threshold-based blocking (e.g., 10 requests per second per IP) is a prudent policy to protect server resources while still allowing the bot to complete its indexing tasks within reasonable timeframes.

Free Traffic Analysis

What's Actually Crawling Your Website?

Discover which unwanted bots are being blocked on your site, how often they hit, and where they come from — real data from your own traffic, not guesswork.

🔍 Scan My Site Free

Powered by JA4 fingerprinting, honeypot traps & behavioral analysis

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.