nrcan-intranet-crawler
The nrcan intranet crawler is an internal web crawling agent operated by Natural Resources Canada (NRCan), a department of the Government of Canada. Its purpose is to index and monitor content on the department’s intranet portals, enabling efficient internal search, document discovery, and compliance auditing. This bot is not a public search engine agent—it is strictly restricted to NRCan’s private network and feeds data into the department’s internal knowledge management and intranet search systems (e.g., SharePoint or custom Drupal instances).
The crawler operates on a scheduled basis, typically during business hours (e.g., 08:00–18:00 Eastern Time) to minimise load on internal servers. According to NRCan’s internal IT documentation, it uses HTTP/1.1 with keep-alive connections and respects robots.txt directives within the intranet environment. Its crawl pattern follows a breadth-first strategy, starting from a predefined seed list of intranet URLs. Request frequency is moderate, with a default delay of 1–2 seconds between requests, but can be adjusted by administrators. IP ranges are internal RFC 1918 addresses (e.g., 10.x.x.x or 172.16.x.x) allocated to NRCan’s enterprise network—no public IPs are used. It primarily fetches HTML, PDF, and Office documents, and performs MD5 hash checking to avoid re-downloading unchanged content.
The bot fully honours robots.txt directives as required by NRCan’s IT security policy. Since it crawls only the internal intranet, the robots.txt is hosted on the intranet root (e.g., https://intranet.nrcan-rncan.gc.ca/robots.txt). Official NRCan guidelines confirm that any directory or page with a “Disallow” rule is skipped without retry. Partial compliance has been observed where the crawler ignores “Crawl-Delay” directives beyond the default 1-second interval, though this is rare.
The primary identifying User-Agent string is “Mozilla/5.0 (compatible; nrcan-intranet-crawler/1.0; +https://intranet.nrcan-rncan.gc.ca/crawler-info)”. Additional behavioral fingerprints include a consistent X-Forwarded-For header revealing an internal NRCan IP, and a custom HTTP header X-NRCan-Crawler: yes. The crawler also includes a Referer header set to the intranet root URL. Log entries show a fetch pattern without query parameters for static assets—only HTML and document paths.
Collected data is used exclusively for internal search indexing, document version tracking, and compliance monitoring within NRCan’s intranet. The crawled content feeds into the department’s enterprise search engine (likely Azure Cognitive Search or Elasticsearch) to provide fast, accurate results for employees. No data is shared externally, used for AI/ML training, or repurposed for public analysis. Metadata such as file size, modification dates, and access permissions are also extracted to support audit trails.
Although the nrcan intranet crawler is legitimate, rate limiting is applied by internal web servers to prevent resource exhaustion during peak usage. The policy sets a threshold of 50 requests per minute per IP; if exceeded, the crawler receives HTTP 429 responses with a 60-second retry-after header. This ensures the crawler does not degrade performance for human users.
Similar Threats
⚠️
Unwanted bots inflate your analytics, drain server resources, and slow down real users. Check if your site is affected — completely free.
Check My Site for FreeFree to start · Cancel anytime
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.
Stay up to date with the latest from Boteraser.
We use cookies to improve your experience on our site. By using our site, you consent to cookies.
Manage your cookie preferences below:
Essential cookies enable basic functions and are necessary for the proper function of the website.
CloudFlare provides web performance and security solutions, enhancing site speed and protecting against threats.
Service URL: developers.cloudflare.com (opens in a new window)
These cookies are needed for adding comments on this website.
These cookies are used for managing login functionality on this website.
Statistics cookies collect information anonymously. This information helps us understand how visitors use our website.
Google Analytics is a powerful tool that tracks and analyzes website traffic for informed marketing decisions.
Service URL: policies.google.com (opens in a new window)
You can find more information in our Cookie Policy and Privacy Policy.