firstgov-gov-search
FirstGov.gov Search is a web crawler operated by the U.S. General Services Administration (GSA), serving as the indexing engine for the official U.S. government web portal originally known as FirstGov.gov (now USA.gov). Its primary purpose is to collect publicly accessible content from federal, state, and local government websites to power the search functionality on USA.gov and other affiliated government search portals. According to GSA documentation, the bot is part of the USA.gov Search service, which replaced the original FirstGov search engine in 2011, but the legacy user-agent string persists. The bot is entirely legitimate and authorized for public-sector information retrieval.
The crawler exhibits moderate crawl patterns, typically fetching pages at a rate of one request every few seconds to avoid overwhelming government servers. It follows standard HTTP/1.1 protocols and respects Cache-Control headers. The IP ranges used are allocated to the GSA’s cloud infrastructure, primarily within the 155.100.0.0/16 and 192.168.0.0/16 blocks for internal testing, but publicly it is known to originate from 13.64.0.0/11 (Azure US Government cloud) based on operational reports. The crawler primarily targets .gov, .mil, and .us domains, but may also index state and local government sites with explicit permission. It does not follow JavaScript-rendered content and focuses on static HTML, sitemaps, and RSS feeds.
The GSA clearly states that the FirstGov.gov Search bot honors Disallow directives in robots.txt files. Official guidance on USA.gov’s developer portal advises webmasters to use standard robots.txt syntax to exclude sections like /admin/ or /private/. There are no documented instances of the bot ignoring crawl-delay or disallow rules; it is considered well-behaved by government IT administrators.
The primary user-agent string is FirstGov.gov Search (also seen as USA.gov Search or usasearch). It does not include a version number. The bot identifies itself via the User-Agent header and may also send a From header with the email [email protected]. Behavioral fingerprints include a consistent crawl interval of 2–5 seconds, no support for cookies, and a user-agent that always begins with “FirstGov”.
Collected data is used exclusively to build the search index for USA.gov and its partner sites, allowing citizens to find government services, forms, and information. The index is updated daily and does not feed any commercial AI training or analytics platforms. The GSA publishes a public search API that returns indexed results, ensuring transparency.
Although the bot is legitimate, it is rate-limited because it can become aggressive during full re-indexing cycles, generating thousands of requests per hour. Threshold-based blocking (e.g., 10 requests per second per IP) is a prudent policy to protect server resources while still allowing the bot to complete its indexing tasks within reasonable timeframes.
Free Traffic Analysis
Discover which unwanted bots are being blocked on your site, how often they hit, and where they come from — real data from your own traffic, not guesswork.
🔍 Scan My Site FreePowered by JA4 fingerprinting, honeypot traps & behavioral analysis
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.
Stay up to date with the latest from Boteraser.
We use cookies to improve your experience on our site. By using our site, you consent to cookies.
Manage your cookie preferences below:
Essential cookies enable basic functions and are necessary for the proper function of the website.
CloudFlare provides web performance and security solutions, enhancing site speed and protecting against threats.
Service URL: developers.cloudflare.com (opens in a new window)
These cookies are needed for adding comments on this website.
These cookies are used for managing login functionality on this website.
Statistics cookies collect information anonymously. This information helps us understand how visitors use our website.
Google Analytics is a powerful tool that tracks and analyzes website traffic for informed marketing decisions.
Service URL: policies.google.com (opens in a new window)
You can find more information in our Cookie Policy and Privacy Policy.