Skip to main content

Boteraser | Website and Server Security Solutions

tapuzbot

Bot User-Agent: tapuzbot

🤖 Overview

Tapuzbot is a legitimate web crawler operated by Tapuz, an Israeli community and portal website (tapuz.co.il), designed to index user-generated content, forums, and articles for internal search and content discovery within the Tapuz platform. According to publicly available records, Tapuzbot has been active since at least 2005 and is used exclusively to feed data into Tapuz’s own search engine and recommendation system, not for third-party AI training. The bot is primarily focused on Hebrew-language websites but can crawl any publicly accessible content.

🌐 Technical Behavior

Tapuzbot follows a classical breadth-first crawl pattern, typically requesting a single page per host every 5–10 seconds under normal operation, though it may burst to 2–3 requests per second during initial site discovery. It uses HTTP/1.1 and supports gzip compression. The IP ranges used by Tapuzbot are largely concentrated within Israeli data centers, notably from Bezeq International (AS8551) and XNET (AS48145), as well as some ranges from Amazon Web Services (EU-West-1) for cloud-based crawling. The bot does not execute JavaScript and only fetches static HTML, CSS, and images necessary for indexing. It respects the If-Modified-Since header to reduce redundant downloads and caches content for up to 24 hours before rechecking.

📋 robots.txt Compliance

Tapuzbot officially honors the Robots Exclusion Standard as documented on Tapuz’s own support pages. It will not crawl any URL disallowed by Disallow directives in robots.txt and also respects Crawl-Delay directives when present. However, community reports from webmasters (e.g., on Israeli hosting forums) indicate that the bot occasionally ignores Disallow for paths that end in ?page= parameters, though Tapuz has since patched this behavior in 2019. Overall, its compliance is considered above average compared to other regional bots.

🔍 Detection Indicators

The primary User-Agent string is TapuzBot/1.0 (compatible; +http://www.tapuz.co.il/bot/), with some variations using Tapuz bot or TapuzBot-Image for image fetching. Behavioral fingerprints include a From header set to [email protected] and a Referer header that always starts with http://www.tapuz.co.il/. The bot also sends a X-Tapuz-Crawl-ID header containing a unique 32-character hex value for tracking purposes.

📊 Data Usage

Collected data is used exclusively for Tapuz’s internal search index and content recommendation engine, which surfaces forum threads, articles, and user profiles on the Tapuz portal. The bot does not contribute to any third-party AI model training or large language model development. Tapuz states in its privacy policy that cached copies are retained for a maximum of 30 days and are not shared with external entities.

⚙️ Rate Limiting Policy

While Tapuzbot is legitimate, it can generate high volumes of requests during initial deep crawls, especially on sites with thousands of forum threads. Rate-limiting is recommended to prevent server resource exhaustion—most webmasters implement a threshold of 20 requests per 10 seconds before blocking, as the bot typically does not require more than that for regular reindexing. The policy rationale is to maintain site stability without permanently banning a benign crawler that supports local content discovery.

🛡️

Stop Bots. Save Bandwidth. Protect Revenue.

Boteraser automatically detects and blocks unwanted bots — protecting your site from scrapers, DDoS bursts, and credential stuffing attacks without slowing down real visitors.

✅ Start Free Protection

Setup takes under a minute  ·  Free trial available

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.