searchpreview
SearchPreview is a legitimate web crawler operated by SearchPreview Ltd. (searchpreview.com), designed to generate thumbnail previews of web pages for use when links are shared on social media, messaging apps, and search result snippets. First documented in 2015, its core function is to capture full-page screenshots using a headless browser, providing visual context without storing textual content for AI training or indexing. The bot operates under a public robots.txt policy and is widely recognized by webmasters.
The crawler uses a headless Chromium engine to render JavaScript and CSS, simulating a real browser. It issues HTTP GET requests at a rate of 10–15 requests per second per IP, sourced from a pool of over 200 IPs within ASN AS16509 (Amazon Web Services) and AS40676 (Psychz Networks) as listed in official documentation. Request timeout is 30 seconds, and the bot follows up to one redirect per URL, stopping at crawl depth 3. It respects gzip encoding and does not crawl binary files or deep nested directories.
According to searchpreview.com/robots.txt, the bot fully honors Disallow directives and supports a configurable Crawl-Delay. It rechecks robots.txt every 24 hours per domain. Webmaster reports confirm immediate cessation on blocked paths, and the bot does not override user-agent exclusions.
Primary User-Agent: SearchPreview/2.0 (compatible; +http://searchpreview.com/bot). Also seen as SearchPreview/1.0. The bot sets an X-Bot header to "searchpreview". Behavioral fingerprint: consistent 2-second inter-request delay, preference for HTML over images or PDFs, and no concurrent requests to the same domain.
Collected data is used solely to generate static preview images (PNG/JPEG) stored for up to 7 days. No textual content is retained, analyzed for AI training, or used for search rankings. The previews are served exclusively via the searchpreview.com CDN to enhance link sharing across platforms like Twitter, Slack, and Facebook.
Rate limiting is applied to prevent server overload; the bot recommends a crawl delay of 5 seconds for small sites and self‑throttles when encountering a 429 status code. It is safe to allow with threshold‑based blocking only if requests exceed 30 per minute to a single domain per its own operational guidelines.
Similar Threats
🛡️
Boteraser automatically detects and blocks unwanted bots — protecting your site from scrapers, DDoS bursts, and credential stuffing attacks without slowing down real visitors.
✅ Start Free ProtectionSetup takes under a minute · Free trial available
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.
Stay up to date with the latest from Boteraser.
We use cookies to improve your experience on our site. By using our site, you consent to cookies.
Manage your cookie preferences below:
Essential cookies enable basic functions and are necessary for the proper function of the website.
CloudFlare provides web performance and security solutions, enhancing site speed and protecting against threats.
Service URL: developers.cloudflare.com (opens in a new window)
These cookies are needed for adding comments on this website.
These cookies are used for managing login functionality on this website.
Statistics cookies collect information anonymously. This information helps us understand how visitors use our website.
Google Analytics is a powerful tool that tracks and analyzes website traffic for informed marketing decisions.
Service URL: policies.google.com (opens in a new window)
You can find more information in our Cookie Policy and Privacy Policy.