tapuzbot
Tapuzbot is a legitimate web crawler operated by Tapuz, an Israeli community and portal website (tapuz.co.il), designed to index user-generated content, forums, and articles for internal search and content discovery within the Tapuz platform. According to publicly available records, Tapuzbot has been active since at least 2005 and is used exclusively to feed data into Tapuz’s own search engine and recommendation system, not for third-party AI training. The bot is primarily focused on Hebrew-language websites but can crawl any publicly accessible content.
Tapuzbot follows a classical breadth-first crawl pattern, typically requesting a single page per host every 5–10 seconds under normal operation, though it may burst to 2–3 requests per second during initial site discovery. It uses HTTP/1.1 and supports gzip compression. The IP ranges used by Tapuzbot are largely concentrated within Israeli data centers, notably from Bezeq International (AS8551) and XNET (AS48145), as well as some ranges from Amazon Web Services (EU-West-1) for cloud-based crawling. The bot does not execute JavaScript and only fetches static HTML, CSS, and images necessary for indexing. It respects the If-Modified-Since header to reduce redundant downloads and caches content for up to 24 hours before rechecking.
Tapuzbot officially honors the Robots Exclusion Standard as documented on Tapuz’s own support pages. It will not crawl any URL disallowed by Disallow directives in robots.txt and also respects Crawl-Delay directives when present. However, community reports from webmasters (e.g., on Israeli hosting forums) indicate that the bot occasionally ignores Disallow for paths that end in ?page= parameters, though Tapuz has since patched this behavior in 2019. Overall, its compliance is considered above average compared to other regional bots.
The primary User-Agent string is TapuzBot/1.0 (compatible; +http://www.tapuz.co.il/bot/), with some variations using Tapuz bot or TapuzBot-Image for image fetching. Behavioral fingerprints include a From header set to [email protected] and a Referer header that always starts with http://www.tapuz.co.il/. The bot also sends a X-Tapuz-Crawl-ID header containing a unique 32-character hex value for tracking purposes.
Collected data is used exclusively for Tapuz’s internal search index and content recommendation engine, which surfaces forum threads, articles, and user profiles on the Tapuz portal. The bot does not contribute to any third-party AI model training or large language model development. Tapuz states in its privacy policy that cached copies are retained for a maximum of 30 days and are not shared with external entities.
While Tapuzbot is legitimate, it can generate high volumes of requests during initial deep crawls, especially on sites with thousands of forum threads. Rate-limiting is recommended to prevent server resource exhaustion—most webmasters implement a threshold of 20 requests per 10 seconds before blocking, as the bot typically does not require more than that for regular reindexing. The policy rationale is to maintain site stability without permanently banning a benign crawler that supports local content discovery.
Similar Threats
🛡️
Boteraser automatically detects and blocks unwanted bots — protecting your site from scrapers, DDoS bursts, and credential stuffing attacks without slowing down real visitors.
✅ Start Free ProtectionSetup takes under a minute · Free trial available
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.
Stay up to date with the latest from Boteraser.
We use cookies to improve your experience on our site. By using our site, you consent to cookies.
Manage your cookie preferences below:
Essential cookies enable basic functions and are necessary for the proper function of the website.
CloudFlare provides web performance and security solutions, enhancing site speed and protecting against threats.
Service URL: developers.cloudflare.com (opens in a new window)
These cookies are needed for adding comments on this website.
These cookies are used for managing login functionality on this website.
Statistics cookies collect information anonymously. This information helps us understand how visitors use our website.
Google Analytics is a powerful tool that tracks and analyzes website traffic for informed marketing decisions.
Service URL: policies.google.com (opens in a new window)
You can find more information in our Cookie Policy and Privacy Policy.