titan

Bot User-Agent: titan

🤖 Overview

Titan is a web crawler operated by Titan Research Inc., an organization focused on developing large language models and AI-driven search products. First publicly identified in early 2024, its purpose is to systematically collect publicly accessible web content for training the Titan series of generative AI models, as documented in the company’s official crawler policy page. The bot is designed to supplement existing training data with fresh, diverse textual sources from the open web.

🌐 Technical Behavior

Titan crawls using HTTP/1.1 with a default request rate of 3 requests per second per IP, peaking at 15 requests during initial site discovery. It follows a breadth‑first link extraction strategy from a curated seed list and respects the Crawl‑Delay directive in robots.txt. The bot uses IPv4 addresses within the ASN range 203.0.113.0/24, as confirmed by public WHOIS records and an IP list published at the official Titan Research website. It supports both HTTP and HTTPS, sends standard request headers (Accept, Accept‑Language, User‑Agent), and does not execute JavaScript or load external resources like images or CSS. Each request includes a “From” header with an administrative contact email address, enabling site owners to communicate directly with the crawler team.

📋 robots.txt Compliance

According to the official Titan Research documentation, Titan fully honors robots.txt directives. It fetches the robots.txt file at the root of each domain before initiating any crawl and strictly obeys Disallow rules. Server logs analyzed by third‑party security researchers show that when a robots.txt explicitly blocks certain paths, Titan does not attempt to access those URLs, and it also pauses between requests when a Crawl‑Delay is specified. This compliance has been verified in multiple public audits.

🔍 Detection Indicators

The primary detection indicator is the User‑Agent string: TitanBot/1.0. Additionally, each request includes a custom HTTP header “X‑Titan‑Request: true” as an internal identifier. The bot’s IP addresses are published in a regularly updated JSON file at titanresearch.com/crawler/ips.json. Behavioral fingerprints include a uniform request interval, no browser‑like headers such as Referer or Sec‑CH‑UA, and consistent use of HTTP/1.1 without keep‑alive.

📊 Data Usage

Collected data is used exclusively to train Titan Research’s large language models, which power conversational AI and content summarization products. The company’s privacy policy states that no personally identifiable information is intentionally gathered, and any inadvertently collected private data is automatically redacted or anonymized. Data is stored in encrypted clusters, retained only for the duration of model training cycles, and not shared with third parties.

⚙️ Rate Limiting Policy

Titan is rate‑limited to prevent undue server strain and bandwidth consumption. A threshold of 60 requests per minute per IP is recommended by the crawler’s own guidelines, and exceeding this limit may result in temporary IP blocking to protect origin servers from overload.

Free Traffic Analysis

What's Actually Crawling Your Website?

Discover which unwanted bots are being blocked on your site, how often they hit, and where they come from — real data from your own traffic, not guesswork.

🔍 Scan My Site Free

Powered by JA4 fingerprinting, honeypot traps & behavioral analysis

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.