search publisher
Search Engine User-Agent:search-publisher
🤖 Overview
Search Publisher is a legitimate web crawler operated by Search Publisher Ltd., a UK-based company that provides content discovery and analytics services for online publishers. First publicly documented in 2019, its primary purpose is to crawl publicly accessible web pages to index articles, blog posts, and news content for the company’s Search Publisher platform—a search engine and analytics dashboard used by publishers to understand how their content is being discovered and linked. The bot feeds data into a proprietary index that powers real-time search results and traffic attribution reports for partner sites.
🌐 Technical Behavior
The crawler follows a breadth-first traversal pattern, starting from a set of seed URLs typically provided by publisher partners or discovered via sitemaps. It submits requests over HTTP/1.1 and HTTP/2 with a concurrency of 10–20 simultaneous connections per domain to avoid overwhelming servers. The default crawl frequency is once every 24 hours per page, but high-authority domains may be rechecked every 6 hours. IP addresses are sourced from Amazon Web Services (AWS) EC2 instances in the us-east-1 and eu-west-1 regions, with ranges documented in the official AWS IP address list (e.g., 3.64.0.0/12, 52.48.0.0/14). The bot respects the If-Modified-Since header and uses conditional GET requests when possible.
📋 robots.txt Compliance
According to the official documentation on the Search Publisher support site (searchpublisher.com/robots-guide), the bot fully honors robots.txt directives, including Disallow, Allow, Crawl-delay, and User-agent matching. The crawler reads the robots.txt file once per domain per session and caches it for 24 hours. There are no known reports of the bot ignoring disallowed paths; however, it does not support wildcards in Disallow directives beyond the standard asterisk.
🔍 Detection Indicators
The primary User-Agent string is SearchPublisher/1.0 (+https://searchpublisher.com/bot). A secondary string SearchPublisher-Mobile/1.0 is used for mobile rendering. The bot always sets the From header to [email protected] and includes a X-SearchPublisher-ID header with a unique crawl session UUID. Reverse DNS lookups on its IPs resolve to ec2-*-*-*-*.compute-1.amazonaws.com. No other identifying headers like Accept-Language are sent.
📊 Data Usage
Collected content—including page title, meta description, headings, body text, and linked anchor text—is used exclusively to populate the Search Publisher search index and to generate anonymized traffic statistics for publisher partners. According to their privacy policy (searchpublisher.com/privacy), the data is never used for AI model training or resold to third parties. Crawled pages are stored for up to 90 days and then expired unless the publisher has an active subscription.
⚙️ Rate Limiting Policy
This bot is rate-limited because it can issue a burst of up to 20 concurrent requests per second from AWS IPs, which may exceed typical threshold-based policies configured for unknown crawlers. Web administrators are advised to set a Crawl-delay: 10 in robots.txt to reduce load, or block the bot’s IP ranges if the crawl volume remains above acceptable limits after notifying Search Publisher support.
53% of Web Traffic Is Bots in 2026
— Imperva Bad Bot Report 2026
How much of your traffic is automated? Get your personal bot traffic report and see exactly what's hitting your server — completely free.
📊 Get My Bot ReportSign up in seconds · No card required
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.