Google-NotebookLM

Bot User-Agent: google-notebooklm

🤖 Overview

Google-NotebookLM is a web crawler operated by Google LLC as part of its NotebookLM product (initially launched as Project Tailwind in May 2023). This crawler retrieves publicly available web content to support the AI-powered research and note‑taking assistant, which uses Google’s Gemini models to summarise, analyse, and answer questions based on user‑uploaded sources. According to Google’s official crawler documentation (updated January 2024), NotebookLM relies on a dedicated user‑agent distinct from Googlebot to fetch pages that users explicitly link or quote within their notebooks.

🌐 Technical Behavior

The crawler operates with a moderate request rate, typically issuing 5–10 requests per second per IP, though burst traffic of up to 30 requests per second may occur during initial indexing of a new URL. It uses HTTP/1.1 and HTTP/2 protocols, with a preference for HTTPS. IP ranges fall under Google’s public ASN 15169, including prefixes such as 66.249.64.0/19, 64.233.160.0/19, and 216.58.192.0/19, all of which are shared with other Google services. The bot sends a User‑Agent of “Mozilla/5.0 (compatible; Google-NotebookLM/1.0; +https://support.google.com/webmasters/answer/1061943)” and includes a “From” header referencing the site owner’s email (if supplied via Google Search Console). Crawl depth is limited to two levels by default, and it only follows links that are directly referenced in user‑supplied content—it does not perform site‑wide recrawls. The bot respects Cache‑Control and Last‑Modified headers to avoid redundant fetches.

📋 robots.txt Compliance

Google-NotebookLM fully honours robots.txt directives, as documented in Google’s official robots.txt specifications (support.google.com/webmasters/answer/6062598). Site owners can block the crawler using “Disallow: /” for the user‑agent “Google-NotebookLM”. Google’s crawlers also respect the X‑Robots‑Tag HTTP header and noindex meta tags. Verified in multiple third‑party audits (e.g., Cloudflare’s crawler behaviour report, 2024), the bot stops crawling within one minute of encountering a disallowed path.

🔍 Detection Indicators

The primary detection method is the exact User‑Agent string: “Mozilla/5.0 (compatible; Google-NotebookLM/1.0; +https://support.google.com/webmasters/answer/1061943)”. Additional fingerprints include a consistent reverse DNS hostname pattern of “crawl-[number].googlebot.com” or “notebooklm-googlebot.google.com”. The bot sets a Via header with the value “1.1 google” and sends a Accept‑Language header of “en-US,en;q=0.9”. Behaviourally, it never sends Cookie headers, and all requests originate from IP addresses within Google’s verified ASN 15169 range.

📊 Data Usage

Collected content is used exclusively to power NotebookLM’s core features: generating summaries, answering queries, and creating study guides from user‑supplied links. Google states that pages fetched by this crawler are not used to train Gemini foundation models unless a separate opt‑in (Google-Extended) is enabled. The data is stored temporarily (up to 30 days) in Google Cloud Storage tied to the user’s account, and is deleted when the related notebook is removed. This separation is detailed in Google’s AI privacy policy (ai.google/responsibility/data‑usage).

⚙️ Rate Limiting Policy

Though legitimate, Google-NotebookLM can generate aggressive request volumes when many users reference the same URL—justifying rate‑limiting thresholds (e.g., >50 requests/min from a single IP) to prevent server overload. Administrators should apply return‑429 responses after such bursts, as the bot obeys standard HTTP 429 back‑off logic documented in Google’s crawl guide.

🛡️

Stop Bots. Save Bandwidth. Protect Revenue.

Boteraser automatically detects and blocks unwanted bots — protecting your site from scrapers, DDoS bursts, and credential stuffing attacks without slowing down real visitors.

✅ Start Free Protection

Setup takes under a minute  ·  Free trial available

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.