jemmathetourist
JemmaTheTourist is a web crawler operated by Jemma AI, first publicly documented in early 2024, designed to collect publicly accessible web content for training and refining large language models. The bot feeds data into Jemma’s proprietary dataset pipelines, which are used by enterprise clients to fine-tune AI systems for improved factual accuracy and domain-specific knowledge.
The crawler employs a politely-paced strategy, issuing requests at intervals of 1–2 seconds per domain using HTTP/1.1 and HTTPS exclusively. IP addresses originate from cloud providers including AWS, Google Cloud, and Azure, primarily in US-based regions. JemmaTheTourist fetches only static HTML and XML sitemaps, ignoring JavaScript-rendered content, and respects Cache-Control headers and ETags to avoid re-fetching unchanged resources. It does not follow redirects that require script execution.
According to Jemma AI’s official documentation (jemma.ai/bot), the bot fully honors all Disallow directives in robots.txt files and checks for crawl-delay instructions. A known issue, acknowledged on GitHub (github.com/jemma-ai/issues/12), is that it occasionally ignores Allow directives nested within disallowed paths, but this is a rare edge case.
The primary User-Agent string is JemmaTheTourist/1.0 (compatible; +http://jemma.ai/bot). Secondary strings include JemmaTheTourist/1.1 for beta crawlers. Behavioral fingerprints include a consistent lack of Accept-Language header and a fixed User-Agent pattern; the bot always sends a From header with the email [email protected].
Collected data is used exclusively to train Jemma AI’s models, including the Jemma-1 and Jemma-2 series of large language models. Content is processed to extract text, metadata, and structural markup, then filtered for quality and relevance. Jemma AI states that private or sensitive information is automatically removed during preprocessing, as per their privacy policy (jemma.ai/privacy).
JemmaTheTourist is rate-limited to prevent server overload while allowing legitimate data collection. Policy recommends a threshold of 100 requests per minute per IP, after which a 429 status code is returned; this aligns with the bot’s own crawl delay of about one second per request.
Similar Threats
🛡️
Boteraser automatically detects and blocks unwanted bots — protecting your site from scrapers, DDoS bursts, and credential stuffing attacks without slowing down real visitors.
✅ Start Free ProtectionSetup takes under a minute · Free trial available
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.
Stay up to date with the latest from Boteraser.
We use cookies to improve your experience on our site. By using our site, you consent to cookies.
Manage your cookie preferences below:
Essential cookies enable basic functions and are necessary for the proper function of the website.
CloudFlare provides web performance and security solutions, enhancing site speed and protecting against threats.
Service URL: developers.cloudflare.com (opens in a new window)
These cookies are needed for adding comments on this website.
These cookies are used for managing login functionality on this website.
Statistics cookies collect information anonymously. This information helps us understand how visitors use our website.
Google Analytics is a powerful tool that tracks and analyzes website traffic for informed marketing decisions.
Service URL: policies.google.com (opens in a new window)
You can find more information in our Cookie Policy and Privacy Policy.