oncrawl
OnCrawl is a legitimate, automated SEO crawler operated by the French company OnCrawl SAS, headquartered in Paris. First released in 2016, it is a professional web crawling tool designed to help website owners, SEO agencies, and digital marketers perform technical SEO audits, monitor site health, and uncover crawl budget inefficiencies. Unlike search engine bots, OnCrawl feeds its data into its own analytics platform, where users can visualize site structure, identify broken links, detect duplicate content, and optimize metadata — it is not used for AI model training or public indexing. The crawler is widely adopted by enterprises and SEO professionals and is listed in the official robots.txt database at robotstxt.org as a known bot (robotstxt.org/db/on-crawl.html).
OnCrawl operates as a configurable, multi-threaded web crawler that can emulate desktop and mobile user agents. By default, it respects the crawl-delay directive and can be set to crawl at speeds ranging from 1 to 50 requests per second (configurable per project). The bot identifies itself with the User-Agent string OnCrawl/1.0 (and variants like OnCrawlBot) and typically resolves from IP ranges owned by OVHcloud, Scaleway, and other French data centers. According to OnCrawl’s official documentation (help.oncrawl.com), the crawler supports HTTP/1.1 and HTTPS, follows redirects up to a configurable limit, and can parse JavaScript-rendered content using a headless Chromium engine. It also sends custom headers including X-OnCrawl-Request: 1 and X-Forwarded-For in certain deployments. Verified logs from webmasters report typical request bursts of 10–30 requests per minute per project, though this can be accelerated for large sites.
OnCrawl fully honors robots.txt directives by default. Its official documentation explicitly states that it checks the Disallow rules for the user agent OnCrawl before crawling each URL. If no specific OnCrawl rules are present, it falls back on the wildcard (*) rules. However, OnCrawl also allows users to override robots.txt for their own projects (e.g., when performing audits on their own domain), but this is configurable per crawl job. The crawler includes a robots.txt validator in its dashboard. Source: help.oncrawl.com/en/articles/1507450-crawler-robots-txt.
Primary detection is via the User-Agent string OnCrawl/1.0 (occasionally OnCrawlBot/1.0). Additional fingerprints include the X-OnCrawl-Request: 1 header, a consistent Accept header of text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8, and the Referer header set to https://www.oncrawl.com on initial requests. Behavioral patterns show sequential crawling of sitemaps (if linked in robots.txt) followed by internal links. Response to X-Robots-Tag noindex directives is also honored.
The data collected by OnCrawl is used exclusively within the OnCrawl SaaS platform for technical SEO reporting and analysis. It generates dashboards showing crawl coverage, page performance, indexation status, and site structure. The data is not sold to third parties, used for AI training, or aggregated across customers — each project’s data remains private to the account holder. OnCrawl’s privacy policy (oncrawl.com/privacy-policy) confirms that raw crawl logs are retained for up to 90 days and then anonymized.
While OnCrawl is not malicious, it can be aggressive if configured with high concurrency (e.g., 50 requests per second on small sites). The recommended rationale for rate-limiting is to protect server resources: thresholds such as 10 requests per second per IP with a 403 response after 10 seconds of sustained load are standard. OnCrawl itself advises users to set Crawl-Delay: 5 in robots.txt to throttle its bot. Source: help.oncrawl.com/en/articles/1507450-crawler-robots-txt.
— Imperva Bad Bot Report 2026
How much of your traffic is automated? Get your personal bot traffic report and see exactly what's hitting your server — completely free.
📊 Get My Bot ReportSign up in seconds · No card required
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.
Stay up to date with the latest from Boteraser.
We use cookies to improve your experience on our site. By using our site, you consent to cookies.
Manage your cookie preferences below:
Essential cookies enable basic functions and are necessary for the proper function of the website.
CloudFlare provides web performance and security solutions, enhancing site speed and protecting against threats.
Service URL: developers.cloudflare.com (opens in a new window)
These cookies are needed for adding comments on this website.
These cookies are used for managing login functionality on this website.
Statistics cookies collect information anonymously. This information helps us understand how visitors use our website.
Google Analytics is a powerful tool that tracks and analyzes website traffic for informed marketing decisions.
Service URL: policies.google.com (opens in a new window)
You can find more information in our Cookie Policy and Privacy Policy.