amazon-kendra
Amazon Kendra is a fully managed intelligent search service provided by Amazon Web Services (AWS), officially launched in November 2019. Its purpose is to surface relevant information from unstructured data across enterprise repositories, documents, and websites. The product feeds into AWS Kendra’s search indexes, enabling natural language querying for corporate knowledge bases and intranets. As part of data ingestion, Kendra operates a dedicated web crawler that indexes publicly available web content when configured by customers.
The Amazon Kendra web crawler performs HTTP/HTTPS requests using a custom bot agent. According to the AWS documentation (https://docs.aws.amazon.com/kendra/latest/dg/crawler.html), the crawler respects standard crawling protocols and uses the AmazonKendra User-Agent string. It originates from AWS-owned IP address ranges published in the AWS IP Address Ranges JSON (https://docs.aws.amazon.com/general/latest/gr/aws-ip-ranges.html). The crawler’s default crawl rate is governed by per‑source quotas – for web sources, Kendra can ingest up to 200 pages per minute per data source (as per AWS service limits). Crawl depth and frequency are configurable through the Kendra console; by default, it follows robots.txt directives and respects nofollow and noindex tags. The bot supports both IPv4 and IPv6 requests.
Amazon Kendra officially honors robots.txt Disallow directives as documented in the AWS Kendra Developer Guide. The crawler checks the file before each crawl session and will not index URLs or paths explicitly forbidden. This behavior is verified by AWS support documentation and community reports (e.g., AWS Knowledge Center articles). There is no known evidence of the bot ignoring robots.txt rules.
The primary User-Agent string is AmazonKendra (case‑sensitive). Additional identifying headers include X-Amz-User-Agent: aws-kendra and a standard User-Agent: AmazonKendra/1.0 (aws; aws-kendra) variant. The bot also sends a From header (optional) with the customer’s configured email. Behavioral fingerprints include sequential, non‑bursty request patterns and a consistent crawl delay that respects the site’s Crawl-delay directive in robots.txt.
Collected data is used exclusively for building and updating Amazon Kendra search indexes for the customer who configured the crawler. The service indexes the textual content of crawled pages to enable natural language search queries. No data is used for external AI training, advertising, or any purpose outside the customer’s Kendra instance. AWS’s Data Processing Addendum governs data handling; content is encrypted at rest and in transit.
Amazon Kendra’s crawler is rate‑limited to prevent excessive load on source servers; the default maximum is 200 pages per minute per data source. This threshold‑based blocking policy ensures the crawler does not overwhelm smaller sites while still allowing efficient indexing for enterprise‑scale deployments. Administrators can further reduce the rate via Kendra’s console.
Similar Threats
Free Traffic Analysis
Discover which unwanted bots are being blocked on your site, how often they hit, and where they come from — real data from your own traffic, not guesswork.
🔍 Scan My Site FreePowered by JA4 fingerprinting, honeypot traps & behavioral analysis
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.
Stay up to date with the latest from Boteraser.
We use cookies to improve your experience on our site. By using our site, you consent to cookies.
Manage your cookie preferences below:
Essential cookies enable basic functions and are necessary for the proper function of the website.
CloudFlare provides web performance and security solutions, enhancing site speed and protecting against threats.
Service URL: developers.cloudflare.com (opens in a new window)
These cookies are needed for adding comments on this website.
These cookies are used for managing login functionality on this website.
Statistics cookies collect information anonymously. This information helps us understand how visitors use our website.
Google Analytics is a powerful tool that tracks and analyzes website traffic for informed marketing decisions.
Service URL: policies.google.com (opens in a new window)
You can find more information in our Cookie Policy and Privacy Policy.