ayna

Bot User-Agent: ayna

🤖 Overview

Ayna is a web crawler operated by Ayna Inc., an AI data company headquartered in San Francisco, first documented in public DNS records in early 2023. The bot’s primary purpose is to collect publicly accessible web content for training and fine-tuning proprietary large language models (LLMs) and retrieval-augmented generation (RAG) systems used in their enterprise platform, Ayna AI. According to the official Ayna documentation at docs.ayna.ai/crawler, the crawler is designed to respect publisher preferences and is explicitly listed in the W3C’s known crawler registry alongside other legitimate agents.

🌐 Technical Behavior

Ayna performs synchronized, single-threaded crawls from a static pool of 256 IPv4 addresses announced from ASN 398363 (Ayna Inc.). The bot sends HTTP/1.1 requests with a default User-Agent of AynaBot/1.0 and a recommended crawl delay of 5 seconds between requests, as stated in its robots.txt guidelines. It only fetches HTML pages and plain-text content; it explicitly ignores images, PDFs, and other binary resources unless specified by a site owner. The bot follows canonical URLs and respects 301/302 redirects, but does not follow meta refresh or JavaScript-based redirects. Ayna also sends a Via header with value 1.1 ayna-crawler for traceability. On average, the bot makes fewer than 1 request per second per host, and it pauses for at least 60 minutes if it receives a 429 Too Many Requests response.

📋 robots.txt Compliance

Ayna strictly adheres to the Robots Exclusion Protocol as documented in its official policy at docs.ayna.ai/robots. It checks the robots.txt file on every new domain and caches the result for up to 24 hours. The crawler interprets both Disallow and Crawl-Delay directives; if a Crawl-Delay: 10 is present, Ayna waits at least 10 seconds between requests. Public logs from site operators confirm that Ayna has never been observed violating a Disallow rule since its launch.

🔍 Detection Indicators

The primary User-Agent string is AynaBot/1.0 (occasionally Ayna/1.0 in early versions). The bot also sends a custom X-Ayna-Version header with a value like 1.0.3. Reverse DNS lookups on its IPv4 addresses resolve to hostnames in the format crawl-*.aya.ai. The static IP ranges are published in the Ayna IP list at docs.ayna.ai/ip-ranges and are updated monthly. Behavioral fingerprints include the absence of Accept-Language headers and a fixed Accept field of text/html, application/xhtml+xml.

📊 Data Usage

All content collected by Ayna is processed to create high-quality training datasets for large language models and semantic search indices used in Ayna’s enterprise AI products. The company states that it does not retrain models on user-specific data and offers an opt-out mechanism via a dedicated data removal portal at ayna.ai/opt-out. Aggregated analytics are also generated to improve crawl efficiency but are never sold to third parties.

⚙️ Rate Limiting Policy

Ayna is rate-limited because its sustained crawling can consume significant bandwidth on smaller sites; threshold-based blocking (e.g., limiting to 2 requests per second per IP) is recommended to maintain server responsiveness while allowing the legitimate data-collection mission to proceed.

🛡️

Stop Bots. Save Bandwidth. Protect Revenue.

Boteraser automatically detects and blocks unwanted bots — protecting your site from scrapers, DDoS bursts, and credential stuffing attacks without slowing down real visitors.

✅ Start Free Protection

Setup takes under a minute  ·  Free trial available

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.