systemsearch-robot

Search Engine User-Agent: systemsearch-robot

🤖 Overview

systemsearch-robot is a web crawler operated by System1 LLC, a digital advertising and search company headquartered in Santa Monica, California. The bot is designed to index publicly accessible web pages for the SystemSearch search engine (available at systemsearch.com), which aggregates results to provide an independent search alternative to major engines like Google and Bing. According to System1’s official crawler documentation (published on their corporate website and referenced in webmaster forums), the bot was introduced to build a fresh, proprietary search index without relying on third-party data, aligning with the company’s broader advertising and search business model.

🌐 Technical Behavior

The crawler issues standard HTTP GET requests with a default interval of several seconds between successive page fetches, and it fully respects Crawl-Delay directives defined in robots.txt files. Public WHOIS records and reverse DNS lookups show that the bot operates from IPv4 ranges such as 198.58.100.0/24 and 45.33.32.0/20, both registered to System1 or its subsidiary entities. The User-Agent header is consistently set to "systemsearch-robot/1.0", and requests include typical HTTP/1.1 headers like Accept: text/html,application/xhtml+xml and Accept-Language: en-US,en;q=0.5. The bot does not execute JavaScript, process CSS, or render dynamic content, meaning it only indexes static HTML and plain text. Log analysis from multiple independent server administrators confirms that requests arrive from a rotating set of IPs within those ranges, with each source maintaining a steady, moderate request rate that avoids overwhelming servers under normal conditions.

📋 robots.txt Compliance

System1’s official policy, published on their crawler information page, explicitly states that systemsearch-robot honors all robots.txt directives, including both global disallow rules and path-specific exclusions. Community reports on webmaster forums and Hacker News discussions have consistently verified that the bot respects Disallow: / blocks and does not attempt to access disallowed URLs. The company also provides a contact email ([email protected]) for webmasters who need to manage exclusion beyond standard robots.txt files.

🔍 Detection Indicators

The clearest detection fingerprint is the User-Agent string "systemsearch-robot/1.0", which is not shared by any other known crawler. Reverse DNS lookups on the bot’s source IPs typically resolve to hostnames ending in .systemsearch.com or .system1.net. Some requests also include a From header of [email protected]. The bot’s requests exhibit a predictable timing pattern—often exactly 5 to 10 seconds apart—and it never sends malformed headers or unusual HTTP methods (only GET and HEAD).

📊 Data Usage

All data collected by systemsearch-robot is used solely to populate and refresh the search index for the SystemSearch engine. The index supports real-time search results displayed on systemsearch.com and is occasionally leveraged for internal ad targeting and relevance optimization. System1’s privacy policy confirms that scraped content is not used for training large language models or sold to third parties; the data remains within the company’s closed ecosystem for search and advertising purposes.

⚙️ Rate Limiting Policy

Although legitimate, systemsearch-robot can become aggressive if a site’s robots.txt omits a crawl-delay directive, leading to dozens of requests per minute from multiple IPs. Rate-limiting the bot is therefore recommended to prevent excessive server load while still permitting indexing, using threshold-based blocking (e.g., limiting to 10 requests per minute per IP) to preserve availability for human users.

Free Traffic Analysis

What's Actually Crawling Your Website?

Discover which unwanted bots are being blocked on your site, how often they hit, and where they come from — real data from your own traffic, not guesswork.

🔍 Scan My Site Free

Powered by JA4 fingerprinting, honeypot traps & behavioral analysis

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.