a6-indexer
a6-indexer is a web crawler operated by A6 Technologies, a private company based in Beijing, China, that specializes in large-scale web data extraction for artificial intelligence model training. First publicly identified in server logs in January 2022, its explicit purpose is to index publicly accessible web pages to feed into A6's proprietary large language model (LLM) and vertical search engine.
The crawler employs a distributed fleet of headless Chromium browsers running on Amazon Web Services EC2 instances (primarily in us-east-1 and eu-west-1 regions) and Google Cloud Platform (us-central1). It sends HTTP/1.1 and occasionally HTTP/2 requests at a sustained rate of 5 to 10 requests per second per IP address, with a crawl interval of 2 to 5 seconds between pages. The bot executes JavaScript to render dynamic content and follows internal redirects. Known IP ranges fall under Amazon AS16509 and Google AS15169.
According to A6's official crawler policy published at https://a6.ai/crawler-policy, a6-indexer fully honors robots.txt Disallow directives. Third-party tracking repositories on GitHub (e.g., https://github.com/ai-crawler-tracker/observations) document isolated incidents where the bot temporarily disregarded rules due to cached queues, but A6 states they rectify such issues upon notification within 48 hours.
The primary User-Agent string is a6-indexer/1.0 (with a Mozilla/5.0 compatibility token). Secondary strings include a6-indexer/2.0 and A6Bot/1.0. The bot does not send a custom X-Robots-Tag header or special identification headers; it is identified solely by the User-Agent field. No CVE entries exist as the crawler is not malicious.
Collected data is used exclusively for training A6's conversational AI assistant "A6 Assistant" and improving the company's internal web search index. A6 claims compliance with the EU General Data Protection Regulation (GDPR) by anonymizing collected content and offers an opt-out mechanism at https://a6.ai/opt-out.
Because a6-indexer generates high request volumes that can degrade performance on smaller hosting environments, webmasters commonly rate-limit it to 3 requests per second or block it entirely after exceeding 1,000 requests per hour. This threshold-based blocking is a standard protective measure that recognizes the crawler's legitimate purpose while safeguarding server resources.
🛡️
Boteraser automatically detects and blocks unwanted bots — protecting your site from scrapers, DDoS bursts, and credential stuffing attacks without slowing down real visitors.
✅ Start Free ProtectionSetup takes under a minute · Free trial available
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.
Stay up to date with the latest from Boteraser.
We use cookies to improve your experience on our site. By using our site, you consent to cookies.
Manage your cookie preferences below:
Essential cookies enable basic functions and are necessary for the proper function of the website.
CloudFlare provides web performance and security solutions, enhancing site speed and protecting against threats.
Service URL: developers.cloudflare.com (opens in a new window)
These cookies are needed for adding comments on this website.
These cookies are used for managing login functionality on this website.
Statistics cookies collect information anonymously. This information helps us understand how visitors use our website.
Google Analytics is a powerful tool that tracks and analyzes website traffic for informed marketing decisions.
Service URL: policies.google.com (opens in a new window)
You can find more information in our Cookie Policy and Privacy Policy.