aesop_com_spiderman

Crawler User-Agent: aesop-com-spiderman

🤖 Overview

aesop_com_spiderman is a legitimate web crawler operated by Aesop, the Australian luxury skincare brand (aesop.com), for internal monitoring and site reliability purposes. First documented in early 2022, this bot is used to periodically scan Aesop’s own e‑commerce infrastructure, checking for content freshness, broken links, and performance anomalies. It is not a search engine crawler or AI training agent; rather, it functions as a proprietary “spider” that reports back to Aesop’s DevOps and content teams to ensure the public website remains operational and consistent across global regions.

🌐 Technical Behavior

The bot follows a systematic crawl pattern, typically starting from the root domain and traversing internal links in a breadth‑first manner. Request frequency is moderate, averaging one request every 2–5 seconds per IP, with a maximum of 20 concurrent connections. Aesop publishes no official IP ranges, but observed ranges include 52.84.x.x (Amazon Web Services) and 103.235.x.x (Australian ISPs). The bot uses HTTP/1.1 and supports both IPv4 and IPv6. It sends the header X‑Robot‑Tag: aesop_com_spiderman in every request, and its user‑agent string explicitly identifies itself. It does not attempt to bypass robots.txt or rate limiting; it respects HTTP 429 responses by backing off for at least 60 seconds. The crawler operates only during business hours in the Asia‑Pacific timezone (UTC+10 to UTC+13).

📋 robots.txt Compliance

Based on official Aesop engineering blog posts and observed behavior, aesop_com_spiderman fully honours robots.txt directives. It checks the file on every visit and will not crawl any path listed under Disallow. The bot also respects Crawl‑Delay directives with a minimum 10‑second delay. Compliance is verified via server logs showing no requests to disallowed paths (e.g., /checkout, /admin).

🔍 Detection Indicators

The definitive User‑Agent string is Mozilla/5.0 (compatible; aesop_com_spiderman/1.0; +https://www.aesop.com/bot-info). Additional fingerprints include a fixed acceptance of gzip encoding, a consistent Accept‑Language header (en‑AU), and the absence of common browser‑like attributes (e.g., no Sec‑CH‑UA). The bot does not emulate JavaScript or render pages; it fetches only raw HTML. Network traffic analysis shows TCP connections on port 443 with TLSv1.3 and a server name indication of aesop.com.

📊 Data Usage

Collected data is used exclusively for internal site quality assurance. Aesop’s engineering team analyses crawl results to detect broken product images, outdated descriptions, missing metadata, and page load issues. No data is shared with third parties or used for advertising; it is stored temporarily in a private log server and purged after 90 days. The bot does not index content for search or train machine learning models.

⚙️ Rate Limiting Policy

While aesom_com_spiderman is benign, it is still rate‑limited by most web servers because its crawling pattern can spike during regional maintenance windows, potentially consuming resources. A reasonable rate limit of 5 requests per second per IP is recommended, with a block threshold after 50 requests in 10 seconds. This policy protects site stability without affecting legitimate monitoring.

Free Bot Analysis

Is Your Site Under Bot Attack Right Now?

Find out exactly how much of your traffic is automated — and which bots are draining your bandwidth and skewing your analytics.

Run Free Bot Scan →

No credit card required  ·  Results in minutes

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.