DomainSigmaCrawler

Crawler User-Agent: domainsigmacrawler

🤖 Overview

DomainSigmaCrawler is a legitimate web crawler operated by DomainSigma LLC, a domain intelligence and cybersecurity company headquartered in the United States. Its primary purpose is to collect publicly accessible web content, including domain registration data, DNS records, website metadata, and security reputation information, which feeds into DomainSigma’s domain profiling and threat detection platform. The crawler is explicitly designed to support domain risk assessment, brand monitoring, and investigative research for security teams, and it operates under a published policy of transparency and webmaster cooperation.

🌐 Technical Behavior

DomainSigmaCrawler employs a distributed crawling architecture using IP addresses sourced from Amazon Web Services (AWS) and other reputable cloud hosting providers, with ranges publicly listed on their official website (domainsigma.com/crawler). The crawler issues HTTP/1.1 and HTTPS requests at a moderate rate of approximately 1–3 requests per second per IP address, and it obeys the Crawl-Delay directive if specified in robots.txt. It uses a canonical User-Agent string and includes a “From” or “Contact” email header for webmaster queries. The crawler avoids heavy concurrent connections and pauses if it receives HTTP 429 (Too Many Requests) or TCP resets, demonstrating responsible crawl behavior consistent with industry guidelines.

📋 robots.txt Compliance

According to DomainSigma’s official documentation published at domainsigma.com/robots-txt, the crawler strictly honors all robots.txt directives, including Disallow rules and Crawl-Delay settings. Site owners can block the crawler entirely by adding User-agent: DomainSigmaCrawler followed by Disallow: / in their robots.txt file. No known violations or ignoring of directives have been reported in public security advisories or webmaster forums.

🔍 Detection Indicators

The primary User-Agent string is “DomainSigmaCrawler/1.0”, often accompanied by a comment like “(+https://domainsigma.com/crawler)” for verification. The crawler also sets a custom HTTP header X-DomainSigma-Crawler: true (documented in their crawler policy page). IP address ranges are listed in an ASN lookup table on their site, and the crawler’s behavior includes a fixed delay between requests, making it distinguishable from aggressive or malicious scanners.

📊 Data Usage

Collected data is aggregated into DomainSigma’s proprietary intelligence platform, which provides domain reputation scores, WHOIS change detection, SSL certificate analysis, and threat correlation. It is used primarily for cybersecurity threat intelligence, brand protection, and investigative lead generation. No collected content is used for generative AI training or resold as raw data; it is processed into structured risk indicators for enterprise customers.

⚙️ Rate Limiting Policy

DomainSigmaCrawler is rate-limited by website operators to prevent resource exhaustion, as its batch crawling can still generate noticeable load on smaller sites. The policy rationale is to protect server stability while allowing legitimate data collection; thresholds are set at IP-level request counts per minute (e.g., 60 req/min), and clients may block the crawler if it exceeds acceptable limits despite its built-in throttling.

Free Traffic Analysis

What's Actually Crawling Your Website?

Discover which unwanted bots are being blocked on your site, how often they hit, and where they come from — real data from your own traffic, not guesswork.

🔍 Scan My Site Free

Powered by JA4 fingerprinting, honeypot traps & behavioral analysis

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.