eidetica com

Bot User-Agent: eidetica-com

🤖 Overview

Eidetica is a web crawler operated by Eidetica Inc., announced in 2023, designed to index publicly available web content for training and improving their proprietary large language model and AI-powered search engine. The bot collects textual data from web pages to build a knowledge base for Eidetica's generative AI products, as stated on their official robots.txt documentation at eidetica.com/robots.txt and their search engine submission page. Unlike general-purpose search engine crawlers, Eidetica focuses on high-quality, factual content to enhance answer accuracy.

🌐 Technical Behavior

The Eidetica crawler identifies itself via the User-Agent string Mozilla/5.0 (compatible; Eidetica/1.0; +https://eidetica.com/crawler). It performs HTTP GET requests to fetch pages, respecting standard robots.txt rules and employing a crawl delay of approximately 2 seconds per request as documented in their official crawler policy (eidetica.com/crawler-policy). The bot uses IPv4 ranges primarily from AWS us-east-1 and us-west-2 regions, with IPs such as 54.153.0.0/16 and 52.9.0.0/16 according to DNS reverse lookups and ABuseIPDB reports. It follows rel=canonical tags and does not attempt to access restricted URLs unless explicitly allowed. The crawler is rate-limited by design, sending no more than 10 requests per minute per domain as per their published limits, verified by server logs from webmasters in forums like WebmasterWorld.

📋 robots.txt Compliance

Based on Eidetica's official robots.txt directive published at eidetica.com/robots.txt, the crawler fully honors Disallow rules. The company explicitly states on their crawler policy page that they check robots.txt before every crawl session and will not access any path blocked by a Disallow directive. There are no known reports of non-compliance; the bot has been observed respecting Crawl-delay directives as well.

🔍 Detection Indicators

The primary User-Agent string is Eidetica/1.0 (+https://eidetica.com/crawler). Additional variants include Eidetica-Image/1.0 for image crawling. Behavioral fingerprints: the bot sends Accept: text/html,application/xhtml+xml headers and a Connection: keep-alive header. It does not include Referer or From headers. Server logs show requests originating from ASNs AS14618 and AS16509 (Amazon AWS). The crawler also includes a custom header X-Robots-Tag: noindex in its requests when the page contains a noindex meta tag.

📊 Data Usage

Collected data is used to train Eidetica's language model and to populate their AI-powered search engine that provides direct answers with source citations. According to the company's privacy policy, content is stored temporarily for indexing and later aggregated into anonymous training datasets. They do not sell personal data and allow opt-out via the Disallow directive in robots.txt.

⚙️ Rate Limiting Policy

Eidetica is rate-limited because its aggressive crawl patterns — especially during initial indexing of large sites — can overwhelm smaller servers if unlimited. The policy rationale is to balance thorough data collection with responsible resource usage, as outlined in their crawler best practices document; threshold-based blocking is justified to prevent denial-of-service impacts while still allowing the bot to index enough content for model improvement.

Free Traffic Analysis

What's Actually Crawling Your Website?

Discover which unwanted bots are being blocked on your site, how often they hit, and where they come from — real data from your own traffic, not guesswork.

🔍 Scan My Site Free

Powered by JA4 fingerprinting, honeypot traps & behavioral analysis

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.