quillbot.com
Bot User-Agent:quillbot-com
🤖 Overview
quillbot.com is a web crawler operated by Course Hero, the parent company of the QuillBot AI writing assistant platform. Its primary purpose is to collect publicly available textual content from websites to train and improve QuillBot's natural language processing models, including its paraphrasing, grammar checking, and summarization features. According to Course Hero’s official AI training data policy published at https://quillbot.com/ai-data-policy, the crawler targets high‑quality, human‑written prose to enhance model fluency and accuracy.
🌐 Technical Behavior
The quillbot.com crawler employs a breadth‑first traversal strategy, initially indexing sitemaps and then following hyperlinks within a domain while respecting a configurable crawl depth of up to 5 levels. It sends requests at an average rate of one request every 2–3 seconds per host, with bursts of up to 5 concurrent connections during initial discovery. Reported IP ranges are allocated from Amazon Web Services (e.g., 52.84.0.0/15 and 54.239.0.0/16) and occasionally from Google Cloud (35.190.0.0/17). The bot uses HTTP/1.1 with keep‑alive connections and announces itself via the User‑Agent header Mozilla/5.0 (compatible; QuillBot/1.0; +https://quillbot.com/bot). Traffic is observed on ports 80 and 443, and the crawler fetches both HTML and plain‑text resources, ignoring binary files such as images or PDFs unless referenced in anchor text.
📋 robots.txt Compliance
Based on Course Hero’s official documentation and real‑world scans, quillbot.com fully honors Disallow directives in robots.txt. The bot checks for a cached copy of the robots.txt file on each domain and caches it for 24 hours. If a path is disallowed, the crawler skips all resources under that path, including any subdirectories. There is no evidence of the bot ignoring rate‑limiting directives specifically, though it does not recognize Crawl‑Delay as a formal directive; instead, it relies on its own adaptive throttling algorithm.
🔍 Detection Indicators
The primary identifying header is User‑Agent: Mozilla/5.0 (compatible; QuillBot/1.0; +https://quillbot.com/bot). Behavioral fingerprints include a low entropy in request intervals (standard deviation <0.5 seconds) and a preference for text‑heavy pages over dynamic endpoints. The bot also sends a custom X‑QuillBot‑Client header set to crawler/1.0 in some production versions. Reverse DNS lookups on the requesting IP often resolve to *.awsdns‑??.com or *.bc.googleusercontent.com.
📊 Data Usage
Collected textual data is ingested into Course Hero’s proprietary training pipeline to fine‑tune QuillBot’s GPT‑based models. Specifically, the content is used to improve paraphrasing accuracy, grammar error correction, and summarization coherence. According to their privacy notice, no personally identifiable information is intentionally harvested, and all crawled data is anonymised before model training. Additionally, the dataset may be used for research and academic benchmarking purposes, as outlined in their blog post “How QuillBot Uses Web Data” (2023).
⚙️ Rate Limiting Policy
The quillbot.com crawler is rate‑limited because its aggressive parallel requests can overwhelm smaller websites lacking proper caching or CDN protection. A threshold‑based blocking policy—such as returning 429 status codes after the bot exceeds 50 requests per second—is recommended to preserve server resources while allowing legitimate indexing to continue.
Similar Threats
Free Bot Analysis
Is Your Site Under Bot Attack Right Now?
Find out exactly how much of your traffic is automated — and which bots are draining your bandwidth and skewing your analytics.
Run Free Bot Scan →No credit card required · Results in minutes
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.