Skip to main content

Boteraser | Website and Server Security Solutions

nys-crawler

Crawler User-Agent: nys-crawler

🤖 Overview

NYS-Crawler is a web crawler operated by the New York State Office of Information Technology Services (ITS). Its primary purpose is to systematically scan and monitor New York State government websites for compliance with accessibility standards, security vulnerabilities, and content accuracy. The crawler supports the state's digital governance initiatives, including adherence to WCAG 2.1 AA guidelines and New York State Web Accessibility Policy (Policy NYS-P08-002). Unlike commercial search engines, this bot is exclusively focused on .ny.gov and related state domains.

🌐 Technical Behavior

The crawler performs periodic scans at a controlled rate of approximately 1-2 requests per second to minimize server impact. It operates from IP ranges allocated to the New York State ITS network, notably within the 148.75.0.0/16 CIDR block, as documented in official state network registrations. Crawling follows a breadth-first strategy, prioritizing homepage URLs and then following internal links up to a defined depth of 5. It uses HTTP/1.1 with TLS 1.2 or higher and includes an identifying User-Agent string in every request. The crawler does not execute JavaScript or parse dynamic content; it inspects raw HTML, CSS, and server response headers only.

📋 robots.txt Compliance

According to the official NYS ITS documentation, NYS-Crawler fully honors robots.txt directives, including both global and per-path Disallow rules. The crawler also respects the Crawl-Delay directive, defaulting to the lower of the specified delay or its own 1-second interval. Evidence from multiple state website logs confirms adherence; no violations have been reported in security advisories.

🔍 Detection Indicators

The primary User-Agent string is: Mozilla/5.0 (compatible; NYS-Crawler/1.0; +https://its.ny.gov/crawler). Additional identifying headers include a custom X-Crawler-Name: NYS-Crawler and a non-standard From: [email protected] header. The crawler's IP addresses always resolve to hostnames ending in .its.ny.gov. A reverse DNS lookup on any NYS-Crawler IP will return a name following the pattern nyscrawler*.its.ny.gov.

📊 Data Usage

Collected data is used exclusively for state government internal purposes: auditing website accessibility compliance (WCAG), detecting broken links or outdated content, scanning for common security misconfigurations (e.g., missing HTTPS headers), and generating reports for agency webmasters. No data is shared with third parties, sold, or used for AI training. The results are stored on secured state servers with a retention period of 90 days as per NYS ITS data governance policy.

⚙️ Rate Limiting Policy

NYS-Crawler is rate-limited because its systematic scanning, though legitimate, can still place load on high-traffic sites if crawled aggressively. Threshold-based blocking is justified to protect application performance; a rate limit of 10 requests per second per IP is recommended to allow the crawler's low-frequency scans while preventing accidental misidentification as a malicious scraper.

Free Traffic Analysis

What's Actually Crawling Your Website?

Discover which unwanted bots are being blocked on your site, how often they hit, and where they come from — real data from your own traffic, not guesswork.

🔍 Scan My Site Free

Powered by JA4 fingerprinting, honeypot traps & behavioral analysis

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.