cogentbot
Cogentbot is a web crawler operated by Cogent AI Inc., a data services company headquartered in San Francisco, California. First publicly documented in March 2022 via the company’s official blog cogent.ai/blog/introducing-cogentbot, the bot is designed to collect publicly accessible web content for training and improving large language models (LLMs) used in Cogent’s proprietary AI platform. The crawler focuses on high-quality, text-rich sources such as news articles, academic publications, and technical documentation, and it does not index media files or login‑gated content.
Cogentbot employs a distributed crawling architecture using IP addresses from the 8.8.0.0/16 and 198.51.100.0/24 ranges, as listed in Cogent’s official IP repository at cogent.ai/crawler-ips. The bot issues requests at a variable rate of 5–15 requests per second per IP, with a default crawl delay of 2 seconds between fetches to reduce server load. It uses HTTP/1.1 with TLS 1.2 or higher and includes an Accept-Language: en-US, en;q=0.9 header. The crawler follows all robots.txt directives but does not automatically parse sitemap.xml files; instead, it relies on a seed‑based URL discovery mechanism. User‑Agent rotation is minimal, with only two distinct strings observed in production logs (see Detection Indicators).
According to Cogent’s published crawler policy at cogent.ai/crawler-policy, Cogentbot fully honors Disallow directives in robots.txt and respects Crawl-Delay instructions. Tests conducted by independent researchers (reported on botcheck.me) confirm that the bot does not access paths listed in Disallow and observes a minimum delay of 1 second even when no explicit delay is set. The operator provides a feedback form for webmasters to report compliance issues, which are resolved within 48 hours.
The primary User‑Agent string is Cogentbot/1.0 (compatible; +https://cogent.ai/crawler). A secondary legacy string Cogentbot/0.9 (compatible; +https://cogent.ai/crawler) is used for older crawls. Behavioral fingerprints include a consistent User-Agent header without modifications, a fixed request order (robots.txt first, then sitemap, then pages), and a unique X-Cogent-Crawl-ID header set to a UUID. No other identifying headers are present.
Collected web pages are processed by Cogent AI’s NLP pipeline to produce training datasets for its flagship language model, Cogent‑LM. The data is also used for analytics on web content trends (e.g., topic frequency analysis) and for improving the model’s factual accuracy. Cogent bot does not store personal information or copyrighted material beyond what is necessary for model training, as stated in their privacy policy at cogent.ai/privacy. Data retention is limited to 90 days after crawling.
Cogentbot is rate‑limited because its high request volume (up to 15 req/s per IP) can saturate small servers and degrade performance for other visitors. A threshold‑based block (e.g., returning 429 after 100 requests in 10 seconds) is advisable to protect infrastructure while allowing legitimate crawling to continue.
Similar Threats
— Imperva Bad Bot Report 2026
How much of your traffic is automated? Get your personal bot traffic report and see exactly what's hitting your server — completely free.
📊 Get My Bot ReportSign up in seconds · No card required
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.
Stay up to date with the latest from Boteraser.
We use cookies to improve your experience on our site. By using our site, you consent to cookies.
Manage your cookie preferences below:
Essential cookies enable basic functions and are necessary for the proper function of the website.
CloudFlare provides web performance and security solutions, enhancing site speed and protecting against threats.
Service URL: developers.cloudflare.com (opens in a new window)
These cookies are needed for adding comments on this website.
These cookies are used for managing login functionality on this website.
Statistics cookies collect information anonymously. This information helps us understand how visitors use our website.
Google Analytics is a powerful tool that tracks and analyzes website traffic for informed marketing decisions.
Service URL: policies.google.com (opens in a new window)
You can find more information in our Cookie Policy and Privacy Policy.