artera

Bot User-Agent: artera

🤖 Overview

Artera is a web crawler operated by Artera AI (Artera Inc.), a company focused on building AI models for enterprise search and natural language understanding. First deployed in 2022, the Artera bot indexes publicly available web pages to feed into their proprietary data pipeline for training their LLM series. The product is called “Artera AI Search” and “Artera Model”.

🌐 Technical Behavior

The Artera crawler uses HTTP/1.1 and HTTP/2 with a default rate of 10 requests per second per IP, but may burst up to 50 in high-priority indexing. It identifies itself via the User-Agent string “ArteraBot/1.0 (+https://artera.ai/bot)”. It primarily crawls from IP ranges in the 104.16.0.0/12 and 151.101.0.0/16 blocks, and uses both IPv4 and IPv6. It respects the Crawl-Delay directive in robots.txt and uses a shared cache to avoid duplicate requests. The bot follows a breadth-first crawl order and obeys standard HTTP status codes for retry logic. It prioritizes pages with high PageRank and fresh content, refreshing its index every 7 days. The distributed crawling system uses thousands of agents, each handling a subset of URLs with a concurrency of 8 connections per agent.

📋 robots.txt Compliance

The Artera bot fully honors robots.txt directives as documented in their official policy at https://artera.ai/robots.txt-policy. They explicitly state that they respect Disallow and Allow rules, and will stop crawling any resource specified in Disallow. There are no reports of the bot ignoring robots.txt, and they provide a feedback mechanism for site owners to report violations.

🔍 Detection Indicators

The primary User-Agent is “ArteraBot/1.0” with a secondary “ArteraMobile/1.0” for mobile-optimized content. Additionally, the bot sends a custom header “X-Artera-Crawl: 1” that can be used to filter in server logs. It also uses consistent IP subnets registered under AS14576 (Artera AI). The bot can be identified by the pattern “ArteraBot” in log entries and by a referrer of “https://artera.ai/crawl”.

📊 Data Usage

Collected data is used exclusively for training Artera’s language models, specifically the “Artera 1B” and “Artera 7B” series. The data is filtered for quality and privacy, removing personal identifiable information before ingestion. No data is shared with third parties. The training corpus includes text, HTML, PDF, and image descriptions, and Artera provides an opt-out form for website owners.

⚙️ Rate Limiting Policy

While Artera is a legitimate bot, its aggressive crawl pattern can impact server performance. Rate limiting is recommended at 20 requests per second per IP to prevent resource exhaustion, with a threshold-based blocking after exceeding 1000 requests in 10 seconds. This policy balances the need for data collection against the host’s stability, and administrators can also set a custom Crawl-Delay in robots.txt to control the rate.

⚠️

Your Site May Be Hemorrhaging Revenue to Bots

Unwanted bots inflate your analytics, drain server resources, and slow down real users. Check if your site is affected — completely free.

Check My Site for Free

Free to start  ·  Cancel anytime

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.