usasearch

Search Engine User-Agent: usasearch

🤖 Overview

UsaSearch is a legitimate web crawler operated by the U.S. General Services Administration (GSA) as part of the Search.gov service (formerly known as USA Search). It is designed to index content from federal, state, and local government websites, providing a unified public search experience across .gov and .mil domains. According to the official Search.gov documentation (search.gov/developer/crawler.html), the bot is used exclusively for building the government-wide search index and is not employed for any commercial or AI training purposes.

🌐 Technical Behavior

The UsaSearch crawler follows a controlled crawl schedule, typically visiting pages at a moderate rate to avoid overwhelming smaller government sites. Official documentation from the GSA indicates a default crawl frequency of approximately one request every 5–10 seconds per site, with the ability to respect custom rate limits set via robots.txt Crawl-delay directives. The crawler uses the HTTP/1.1 protocol and sends standard GET requests. IP ranges are dynamically assigned from GSA’s cloud infrastructure (primarily AWS and GSA-owned blocks), but the bot consistently identifies itself with the user-agent string "usasearch" (with optional version, e.g., "usasearch/1.0"). The crawler does not follow JavaScript-rendered content or execute client-side scripts; it retrieves only static HTML and associated metadata.

📋 robots.txt Compliance

The UsaSearch bot fully honors robots.txt Disallow and Allow directives, as confirmed by the Search.gov Support page (search.gov/manual/crawling-indexing.html). The GSA explicitly states that webmasters can block the crawler entirely by adding "User-agent: usasearch — Disallow: /" to their robots.txt file. Compliance has been verified by multiple government security audits, with no reports of the bot ignoring exclusion rules.

🔍 Detection Indicators

The primary identifying string is "usasearch" in the User-Agent header — examples include "usasearch/1.0 (compatible; GSA; +https://search.gov/)" or simply "usasearch". No additional custom headers like X-Robot are used, but the crawler consistently includes a Referer header pointing to search.gov. Behavioral fingerprints include a low request rate and a pattern of crawling only during U.S. business hours (ET), as per the GSA’s operational guidelines.

📊 Data Usage

All data collected by the UsaSearch crawler is used exclusively to populate the Search.gov public search index, which powers search boxes on tens of thousands of federal, state, and local government websites. The index is refreshed periodically (typically every 7–30 days depending on site update frequency). No data is sold, shared with third parties, or used for AI model training. The GSA’s privacy policy (search.gov/policy/privacy.html) confirms that crawler data is stored securely and only used for search relevance improvements.

⚙️ Rate Limiting Policy

While UsaSearch is non-malicious and cooperative, it is rate-limited to prevent its moderate crawl load from affecting the performance of resource-constrained government sites. The policy recommends setting a Crawl-delay of at least 10 seconds in robots.txt if the bot’s default pace is too aggressive, and threshold-based blocking is justified only when the bot fails to comply — an extremely rare occurrence documented in only a handful of cases over a decade of operation.

Free Traffic Analysis

What's Actually Crawling Your Website?

Discover which unwanted bots are being blocked on your site, how often they hit, and where they come from — real data from your own traffic, not guesswork.

🔍 Scan My Site Free

Powered by JA4 fingerprinting, honeypot traps & behavioral analysis

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.