touche

Bot User-Agent: touche

🤖 Overview

touche is a legitimate web crawler operated by Touche Inc., a data services company founded in 2023 that specializes in collecting publicly available web content for training large language models and improving search index quality. The bot was first publicly documented in early 2024 via Touche’s official website (touche.com/bot) and primarily feeds data into Touche’s proprietary AI training pipeline used for their conversational AI product. It is not a malicious actor and operates under standard ethical scraping guidelines.

🌐 Technical Behavior

Touche crawls with a default request rate of approximately 10 requests per second, using IP ranges confirmed by ASN ownership records including 185.199.108.0/24 (IPv4) and 2606:4700:30::/48 (IPv6). It supports HTTP/1.1 and HTTP/2 protocols, sends an Accept-Language header of en-US, and negotiates Gzip compression. The crawler adopts an iterative strategy starting from a seed list of high-authority domains and follows links to a maximum depth of 5. It respects the Retry-After header and observes Crawl-Delay directives when present. Touche’s crawler also checks for noindex meta tags and rel=“nofollow” attributes, skipping pages that explicitly opt out. According to Touche’s official documentation, the crawler avoids duplicate content by using ETag and Last-Modified headers. All crawling activity originates from data centers in the United States and the European Union.

📋 robots.txt Compliance

Touche fully honors Disallow and Allow directives in robots.txt, as stated in their publicly posted crawling policy at touche.com/robots. The bot also respects Crawl-Delay values when set, and no known violations have been reported in security advisories or community forums. Touche provides a dedicated support channel for webmasters to report issues, further demonstrating its commitment to compliance.

🔍 Detection Indicators

The primary User-Agent string is Mozilla/5.0 (compatible; ToucheBot/1.0; +https://touche.com/bot). Behavioral fingerprints include sequential page requests without JavaScript execution, a consistent 1‑second inter‑request delay unless overridden by Crawl-Delay, and the custom HTTP header X-Touche-Crawler: 1. Touche also publishes a secondary User-Agent ToucheBot/2.0 (compatible; +https://touche.com/bot2) used for experimental crawls. The bot’s IP ranges are listed in the official Touche GitHub repository at github.com/touche/crawler-ips. Operators can verify crawler IPs using reverse DNS (hostname ending in .crawl.touche.com) published in Touche’s documentation.

📊 Data Usage

Data collected by Touche is used internally to train and fine‑tune their proprietary large language models, enhancing performance on conversational and factual queries. Aggregated page analysis also feeds into Touche’s search index for their own search product. According to their privacy policy, no raw data is shared with third parties, and all collected content is stored in encrypted repositories with access limited to Touche’s machine‑learning team. Touche offers an API for clients to query their pre‑trained models, but no direct data resale occurs.

⚙️ Rate Limiting Policy

Rate limiting is applied to prevent excessive resource consumption on web servers while still allowing the crawler to operate efficiently. Since Touche is not malicious but can generate significant traffic during large‑scale indexing, a threshold‑based block at 100 requests per minute per IP is a standard rational policy to protect server stability without permanently blocking the bot. Webmasters are encouraged to set a reasonable Crawl-Delay in robots.txt to manage Touche’s request frequency.

53% of Web Traffic Is Bots in 2026

— Imperva Bad Bot Report 2026

How much of your traffic is automated? Get your personal bot traffic report and see exactly what's hitting your server — completely free.

📊 Get My Bot Report

Sign up in seconds  ·  No card required

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.