twikle

Bot User-Agent: twikle

🤖 Overview

Twikle is a search engine bot operated by Twikle Search (twikle.com), a privacy-focused search engine that emerged in 2023. It indexes publicly accessible web content to provide organic search results without tracking users or storing personal data. The bot’s primary function is to crawl and index web pages for the Twikle search index, which is built on a custom retrieval system that combines index-based and semantic search approaches.

🌐 Technical Behavior

Twikle crawlers follow standard HTTP/1.1 and HTTPS protocols, using a crawl frequency that varies between once every 24 hours for most pages to multiple times daily for high-importance sites. The bot originates from a narrow IP range documented in official twikle.com documentation: 146.70.0.0–146.70.255.255 (AS396982). It respects Last-Modified and ETag headers to reduce server load during re-crawls. Twikle’s crawler does not request metadata from third-party analytics services and does not execute JavaScript by default, focusing instead on static HTML content. The bot uses a polite crawl delay of 5–10 seconds between requests per domain, as noted in its technical guidelines published at twikle.com/crawler.

📋 robots.txt Compliance

Based on documented evidence from the Twikle Crawler Guide, Twikle fully adheres to standard robots.txt directives including Disallow, Crawl-delay, and User-agent. The bot processes robots.txt at the start of each crawl session and re-checks the file every 24 hours. Failure to honor Disallow directives would cause the bot to be blocked from further crawling, according to Twikle’s operational policy.

🔍 Detection Indicators

The primary User-Agent string is: TwikleCrawler/1.0 (+https://twikle.com/crawler). Behavioral fingerprints include a consistent request interval of 5–10 seconds, no HTTP referrer other than twikle.com, and no Accept-Language header. The bot also includes a custom X-Robots-Tag header (noindex,nofollow) when instructed by site owners. No CVE entries or security advisories have been associated with Twikle because it is a non-aggressive, legitimate crawler.

📊 Data Usage

Collected data — including page titles, meta descriptions, heading structures, and full text content — is used exclusively to build and maintain the Twikle search index. Twikle does not sell or share crawled data for AI training, advertising, or analytics. The company’s privacy policy (twikle.com/privacy) explicitly states that no user data is stored beyond the crawl logs required for operational integrity.

⚙️ Rate Limiting Policy

Rate limiting is applied to prevent excessive load on origin servers and to enforce the polite crawl delay. In most web application firewalls (WAFs), thresholds of 20 requests per minute from the Twikle IP range are sufficient. If Twikle exceeds 20 requests per minute without honoring Crawl-delay, it should be rate-limited rather than blocked, as the bot will self-throttle after receiving HTTP 429 responses.

Free Traffic Analysis

What's Actually Crawling Your Website?

Discover which unwanted bots are being blocked on your site, how often they hit, and where they come from — real data from your own traffic, not guesswork.

🔍 Scan My Site Free

Powered by JA4 fingerprinting, honeypot traps & behavioral analysis

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.