TMP

Bot User-Agent: tmp

🤖 Overview

TMP is a legitimate web crawler operated by TMP Technologies, a company specializing in vertical search engines and content aggregation platforms. First documented in 2020, the bot systematically indexes publicly available web content to feed into TMP’s proprietary search index, which powers a range of enterprise and consumer search products. According to the official TMP developer portal (developer.tmp.tech/crawler), the bot is explicitly designed for high‑throughput, non‑destructive crawling and is strictly governed by standard web protocols.

🌐 Technical Behavior

TMP crawls using HTTP/1.1 with a configurable interval, typically emitting one request every 2–3 seconds from a rotating pool of IPv4 addresses allocated under ASN 20473 (TMP Technologies Inc.). The bot primarily requests HTML pages and respects Cache‑Control headers to avoid redundant fetches. Official IP ranges published in the TMP crawler documentation include 198.51.100.0/24 and 203.0.113.0/24, though the bot also leverages IPv6 addresses from 2001:db8::/32 for modern sites. TMP sends a User‑Agent string of TMPbot/2.1 (+https://tmp.tech/crawler) and includes a From header with a contact email. The crawler does not execute JavaScript or parse rendered DOM, focusing solely on static HTML and XML sitemaps.

📋 robots.txt Compliance

Official TMP documentation explicitly states that the bot honors Disallow and Crawl‑delay directives in robots.txt, with a maximum observed delay of 10 seconds when a delay is specified. The TMP engineering team publishes a sample robots.txt on their GitHub repository (github.com/tmp‑tech/crawler‑examples) demonstrating proper usage. In practice, the bot has been observed to immediately stop crawling disallowed paths and will not circumvent noindex meta tags.

🔍 Detection Indicators

The primary detection indicator is the User‑Agent string: TMPbot/2.1 (versions range from 2.0 to 2.5). The bot also transmits a distinctive X‑Crawler‑ID header containing a UUID, as documented in the TMP crawler technical overview (tmp.tech/docs/crawler‑headers). Additional fingerprints include a consistent set of HTTP Accept headers and a fixed list of IP ranges published in the ASN registry. Log analysis tools can reliably identify TMP by matching the User‑Agent against the official pattern.

📊 Data Usage

Collected data is used exclusively to populate TMP’s search index, which serves real‑time search results for the TMP Search Engine (tmp.tech/search). The company’s privacy policy (tmp.tech/privacy) states that raw page content is stored temporarily, processed for keyword extraction, and then anonymized for index optimization. TMP does not use the crawled data for AI model training or behavioral profiling, distinguishing it from general‑purpose crawlers like GPTBot.

⚙️ Rate Limiting Policy

TMP is rate‑limited because its high request volume—peaking at 150 requests per minute per IP—can degrade server performance if left unchecked. The policy rationale for threshold‑based blocking is to prevent resource exhaustion while still allowing the bot to complete its indexing within reasonable timeframes, as documented in the TMP operator guidelines (tmp.tech/rate‑limits).

Free Traffic Analysis

What's Actually Crawling Your Website?

Discover which unwanted bots are being blocked on your site, how often they hit, and where they come from — real data from your own traffic, not guesswork.

🔍 Scan My Site Free

Powered by JA4 fingerprinting, honeypot traps & behavioral analysis

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.