ApifyWebsiteContentCrawler
Crawler User-Agent:apifywebsitecontentcrawler
🤖 Overview
ApifyWebsiteContentCrawler is a legitimate web crawler operated by Apify Technologies, a Czech company founded in 2015 that provides a cloud-based web scraping and automation platform. This crawler is the default engine used by Apify’s Website Content Crawler actor, which is publicly available on the Apify Store for users to extract structured data from websites for purposes such as AI training, content aggregation, and competitive analysis. According to Apify’s official documentation (docs.apify.com), the crawler is designed to be respectful of website resources and is not intended for malicious activities; it processes publicly accessible pages only.
🌐 Technical Behavior
The ApifyWebsiteContentCrawler operates by sending HTTP(S) GET requests to target URLs, respecting robots.txt by default, and can follow links up to a configurable depth (default 3 levels). Request frequency is adjustable by the user through Apify’s interface, but the platform enforces a default concurrency limit to avoid overwhelming servers. The crawler’s IP ranges come from Apify’s own data center infrastructure, hosted primarily on Amazon Web Services (AWS) and Google Cloud Platform, with IP blocks that vary per region. Apify publishes a list of known egress IP ranges in their API documentation (apify.com/security). The crawler supports both HTTP/1.1 and HTTP/2, and sends a User-Agent header identifying itself as ApifyWebsiteContentCrawler/1.0 (+https://apify.com), along with an optional X-Requested-With: XMLHttpRequest header when emulating AJAX calls.
📋 robots.txt Compliance
According to Apify’s official documentation on robots.txt handling (docs.apify.com/platform/actors/development/apify-library/robots-txt), the ApifyWebsiteContentCrawler fully complies with robots.txt directives by default. The crawler reads the Disallow rules and does not crawl paths explicitly forbidden. However, users can override this behavior in the actor configuration by disabling the “Respect robots.txt” option, which is clearly documented as a setting for advanced users who assume responsibility for compliance.
🔍 Detection Indicators
The primary identifying header is User-Agent: ApifyWebsiteContentCrawler/1.0 (+https://apify.com). For JavaScript-rendered pages, the crawler may also include a X-Apify-Request-Id header used for internal tracking. Behavioral fingerprints include a consistent request interval (default 1 second between pages) and a tendency to request /robots.txt before any other page. Web servers can detect the crawler by monitoring for repeated pattern of GET /robots.txt followed by a burst of page requests within the same IP range.
📊 Data Usage
Data collected by the ApifyWebsiteContentCrawler is stored in the user’s Apify account and used for the specific purpose defined by the actor’s configuration, which may include AI training datasets, search indexing preprocessing, or general analytics. Apify does not claim ownership of scraped data; the data remains under the control of the user who runs the actor. Apify’s terms of service require users to ensure they have legal rights to scrape the target websites.
⚙️ Rate Limiting Policy
Rate limiting of the ApifyWebsiteContentCrawler is recommended because, while the crawler is rate-limited by default, aggressive user configurations (e.g., reducing delay to zero) can cause performance issues on target servers. A threshold-based blocking policy (e.g., 10+ requests per second from a single Apify IP) is a prudent measure to protect server resources while still allowing legitimate crawling activity.
Similar Threats
🛡️
Stop Bots. Save Bandwidth. Protect Revenue.
Boteraser automatically detects and blocks unwanted bots — protecting your site from scrapers, DDoS bursts, and credential stuffing attacks without slowing down real visitors.
✅ Start Free ProtectionSetup takes under a minute · Free trial available
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.