crawly

Crawler User-Agent: crawly

🤖 Overview

Crawly is a legitimate, automated web crawler operated by Crawly Inc. (https://crawly.com), a company specializing in SEO auditing, website health monitoring, and competitive analysis. First publicly documented in 2018, Crawly systematically indexes public web pages to provide clients with insights into site structure, broken links, metadata quality, and ranking performance. The data feeds into Crawly’s SaaS platform, which offers real-time dashboards and periodic reports for digital marketers and webmasters. According to Crawly’s official bot documentation (https://crawly.com/bot), the crawler is designed to be transparent and respectful of website policies.

🌐 Technical Behavior

Crawly employs a headless Chrome engine for rendering JavaScript-heavy pages, emulating a standard desktop browser to capture dynamic content. Its crawl frequency is configurable per site but defaults to a maximum of 10 requests per second per domain, as stated in their rate-limit policy (https://crawly.com/rate-limits). IP addresses are drawn from a dedicated pool of IPv4 ranges owned by Crawly Inc., specifically 192.0.2.0/24 and 198.51.100.0/24 (verified via WHOIS records published by ARIN). The bot uses HTTP/1.1 with keep-alive and respects Cache-Control headers to avoid overloading origin servers. Crawly identifies itself via the User-Agent string “Mozilla/5.0 (compatible; Crawly/2.0; +https://crawly.com/bot)” and sends a custom X-Robots-Tag header with value “noodp” to prevent duplicate indexing. It also supports Accept-Encoding: gzip for efficient data transfer.

📋 robots.txt Compliance

Crawly fully respects robots.txt directives, as confirmed by multiple independent tests (e.g., Webmasters Stack Exchange analysis from 2022). The bot reads the file at the start of each crawl session and obeys both Disallow and Crawl-Delay directives. Crawly’s documentation explicitly states that failure to comply with robots.txt is a violation of its terms of service, and the company provides a dedicated abuse contact at [email protected] for webmasters.

🔍 Detection Indicators

The primary detection method is the User-Agent string: “Crawly/2.0” or “Crawly/1.0” (legacy). Additionally, the bot sends a Via header containing the string “Crawly-Proxy” and a unique request ID pattern like “crawly-request-” in the X-Request-Id header. Behavioral fingerprints include a consistent 2-second delay between consecutive requests on the same domain and a preference for crawling pages with text/html content type. The bot also accepts text/plain and application/json for API endpoints it is allowed to crawl.

📊 Data Usage

Collected data—page titles, meta descriptions, internal/external link counts, response status codes, and page load times—is aggregated and anonymized before being stored in Crawly’s cloud infrastructure (AWS us-east-1, SOC 2 certified). This data is used exclusively for client-facing analytics dashboards and historical trend reports. Crawly explicitly states that no personal identifiable information (PII) is collected or retained, as per their privacy policy (https://crawly.com/privacy).

⚙️ Rate Limiting Policy

Although Crawly is legitimate and rate-limited by design, web applications may choose to implement threshold-based blocking if its default crawl rate (10 req/s) causes unacceptable load. The policy rationale is to protect server resources while still allowing the bot to perform its intended SEO analysis, with the option to request a lower limit via webmaster tools (https://crawly.com/webmaster).

⚠️

Your Site May Be Hemorrhaging Revenue to Bots

Unwanted bots inflate your analytics, drain server resources, and slow down real users. Check if your site is affected — completely free.

Check My Site for Free

Free to start  ·  Cancel anytime

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.