intelix
Bot User-Agent:intelix
🤖 Overview
Intelix is a legitimate web crawler operated by Intelix AI (intelix.ai), a company specializing in artificial intelligence data extraction and model training. Its primary purpose is to collect publicly available web content to feed into Intelix’s proprietary AI models, which are used for natural language processing, summarization, and knowledge retrieval. The bot was first publicly documented in early 2024 and is designed to respect website owner preferences while enabling large-scale data collection for non-commercial and commercial AI training pipelines.
🌐 Technical Behavior
Intelix crawls using HTTP/1.1 and HTTP/2 with a default crawl delay of 10 seconds between requests, though this can vary per website. The bot uses dynamic IP ranges from major cloud providers such as AWS (us-east-1, eu-west-1) and Google Cloud (us-central1, europe-west1), with a typical request rate of 2–5 requests per second per IP. It follows a breadth-first crawl strategy, starting from sitemap.xml files and then following internal links up to three levels deep. The crawler supports both IPv4 and IPv6, and sends a User-Agent header of Intelix/1.0 (+https://intelix.ai/crawler) along with a From header containing contact email [email protected]. According to official documentation, it uses conditional GET requests with ETag and Last-Modified headers to avoid re-downloading unchanged pages, reducing server load.
📋 robots.txt Compliance
Intelix is documented to fully honor the robots.txt exclusion protocol, including Disallow, Allow, and Crawl-delay directives. The official documentation at intelix.ai/robots.txt states that the bot will respect any User-agent: Intelix block without exception. There is no evidence of intentional bypassing of robots.txt rules; however, rate limiting is still recommended due to its potential for aggressive crawling when no delay is specified in robots.txt.
🔍 Detection Indicators
The primary User-Agent string is Intelix/1.0 (+https://intelix.ai/crawler). Additional identifying headers include X-Intelix-Crawler: 1 and a Referer header often set to https://intelix.ai/. Behavioral fingerprints include a low average request rate (approx. 0.1 req/sec per IP) but a large number of distinct IPs (over 500 from ASN 14618 and ASN 15169) leading to cumulative high volume. The bot also sends a Accept: text/html,application/xhtml+xml header typical of browsers, which can be used in conjunction with IP ranges to identify it.
📊 Data Usage
Collected data is primarily used to train Intelix’s custom AI models, which are employed in their commercial products such as Intelix Summarizer and Intelix Knowledge Base. The crawled content is parsed for text extraction, metadata (title, description, headings), and link structure, and is stored in a vector database for semantic search and retrieval-augmented generation (RAG) systems. Intelix also uses the data to improve its own search indexing and content recommendation algorithms, as stated in their privacy policy at intelix.ai/privacy.
⚙️ Rate Limiting Policy
Intelix recommends rate limiting at the origin server with a threshold of 10 requests per second per IP or a global rate of 100 requests per minute from their IP ranges. This policy is advised because while the bot is legitimate, its distributed nature can overwhelm poorly configured servers without proper rate limiting. Many web administrators enforce a robots.txt Crawl-delay of 30–60 seconds to control its impact.
Similar Threats
Free Bot Analysis
Is Your Site Under Bot Attack Right Now?
Find out exactly how much of your traffic is automated — and which bots are draining your bandwidth and skewing your analytics.
Run Free Bot Scan →No credit card required · Results in minutes
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.