searchmarks

Search Engine User-Agent: searchmarks

🤖 Overview

Searchmarks is a web crawler operated by Searchmarks Inc., a meta-search engine launched in 2019 that aggregates results from Google, Bing, and other sources. Its primary purpose is to index publicly accessible web content to fuel the Searchmarks search engine, which processes millions of queries daily as stated on the official site at searchmarks.com.

🌐 Technical Behavior

The crawler uses HTTP/1.1 and HTTPS, making requests at a default rate of one request every five seconds per host, configurable via the Crawl-Delay directive. It employs a distributed architecture with IP addresses from the 45.33.32.0/19 and 69.164.200.0/18 blocks, documented in the official bot information page. It follows hyperlinks recursively up to a depth of 10, respects canonical URLs, and includes an Accept-Language header set to en-US. The bot supports the If-Modified-Since header to reduce bandwidth and sets a custom HTTP header X-Searchmarks-Bot: true for identification. It uses asynchronous requests and does not execute JavaScript, focusing solely on static HTML content.

📋 robots.txt Compliance

Searchmarks fully honors robots.txt Disallow directives and ignores any directories or files marked as disallowed, as confirmed by its official robots.txt compliance page at searchmarks.com/bot. It also respects the Crawl-Delay directive and pauses accordingly, and supports Allow directives for granular control.

🔍 Detection Indicators

Distinct User-Agent strings include Searchmarks/1.0 and Mozilla/5.0 (compatible; Searchmarks/1.0; +http://www.searchmarks.com/bot). Additional fingerprints include the HTTP header From: [email protected] and a consistent request pattern with no JavaScript or cookie usage. The bot’s IP ranges are publicly listed in the 45.33.32.0/19 and 69.164.200.0/18 subnets.

📊 Data Usage

Collected data is used exclusively for the Searchmarks meta-search engine, combining results from multiple providers to give users a unified search experience. No content is sold, used for AI model training, or shared with third parties. Site owners can opt out by adding a <meta name="searchmarks" content="noindex"> tag to their pages.

⚙️ Rate Limiting Policy

Rate limiting is recommended because the crawler can become aggressive during initial indexing of a new site. Threshold-based blocking, such as limiting requests to 10 per second per IP, protects server resources while allowing legitimate crawling activities.

🛡️

Stop Bots. Save Bandwidth. Protect Revenue.

Boteraser automatically detects and blocks unwanted bots — protecting your site from scrapers, DDoS bursts, and credential stuffing attacks without slowing down real visitors.

✅ Start Free Protection

Setup takes under a minute  ·  Free trial available

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.