jemmathetourist
Bot User-Agent:jemmathetourist
🤖 Overview
JemmaTheTourist is a web crawler operated by Jemma AI, first publicly documented in early 2024, designed to collect publicly accessible web content for training and refining large language models. The bot feeds data into Jemma’s proprietary dataset pipelines, which are used by enterprise clients to fine-tune AI systems for improved factual accuracy and domain-specific knowledge.
🌐 Technical Behavior
The crawler employs a politely-paced strategy, issuing requests at intervals of 1–2 seconds per domain using HTTP/1.1 and HTTPS exclusively. IP addresses originate from cloud providers including AWS, Google Cloud, and Azure, primarily in US-based regions. JemmaTheTourist fetches only static HTML and XML sitemaps, ignoring JavaScript-rendered content, and respects Cache-Control headers and ETags to avoid re-fetching unchanged resources. It does not follow redirects that require script execution.
📋 robots.txt Compliance
According to Jemma AI’s official documentation (jemma.ai/bot), the bot fully honors all Disallow directives in robots.txt files and checks for crawl-delay instructions. A known issue, acknowledged on GitHub (github.com/jemma-ai/issues/12), is that it occasionally ignores Allow directives nested within disallowed paths, but this is a rare edge case.
🔍 Detection Indicators
The primary User-Agent string is JemmaTheTourist/1.0 (compatible; +http://jemma.ai/bot). Secondary strings include JemmaTheTourist/1.1 for beta crawlers. Behavioral fingerprints include a consistent lack of Accept-Language header and a fixed User-Agent pattern; the bot always sends a From header with the email [email protected].
📊 Data Usage
Collected data is used exclusively to train Jemma AI’s models, including the Jemma-1 and Jemma-2 series of large language models. Content is processed to extract text, metadata, and structural markup, then filtered for quality and relevance. Jemma AI states that private or sensitive information is automatically removed during preprocessing, as per their privacy policy (jemma.ai/privacy).
⚙️ Rate Limiting Policy
JemmaTheTourist is rate-limited to prevent server overload while allowing legitimate data collection. Policy recommends a threshold of 100 requests per minute per IP, after which a 429 status code is returned; this aligns with the bot’s own crawl delay of about one second per request.
Similar Threats
🛡️
Stop Bots. Save Bandwidth. Protect Revenue.
Boteraser automatically detects and blocks unwanted bots — protecting your site from scrapers, DDoS bursts, and credential stuffing attacks without slowing down real visitors.
✅ Start Free ProtectionSetup takes under a minute · Free trial available
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.