google-notebooklm
Google-NotebookLM is a web crawler operated by Google LLC as part of its NotebookLM product (initially launched as Project Tailwind in May 2023). This crawler retrieves publicly available web content to support the AI-powered research and note‑taking assistant, which uses Google’s Gemini models to summarise, analyse, and answer questions based on user‑uploaded sources. According to Google’s official crawler documentation (updated January 2024), NotebookLM relies on a dedicated user‑agent distinct from Googlebot to fetch pages that users explicitly link or quote within their notebooks.
The crawler operates with a moderate request rate, typically issuing 5–10 requests per second per IP, though burst traffic of up to 30 requests per second may occur during initial indexing of a new URL. It uses HTTP/1.1 and HTTP/2 protocols, with a preference for HTTPS. IP ranges fall under Google’s public ASN 15169, including prefixes such as 66.249.64.0/19, 64.233.160.0/19, and 216.58.192.0/19, all of which are shared with other Google services. The bot sends a User‑Agent of “Mozilla/5.0 (compatible; Google-NotebookLM/1.0; +https://support.google.com/webmasters/answer/1061943)” and includes a “From” header referencing the site owner’s email (if supplied via Google Search Console). Crawl depth is limited to two levels by default, and it only follows links that are directly referenced in user‑supplied content—it does not perform site‑wide recrawls. The bot respects Cache‑Control and Last‑Modified headers to avoid redundant fetches.
Google-NotebookLM fully honours robots.txt directives, as documented in Google’s official robots.txt specifications (support.google.com/webmasters/answer/6062598). Site owners can block the crawler using “Disallow: /” for the user‑agent “Google-NotebookLM”. Google’s crawlers also respect the X‑Robots‑Tag HTTP header and noindex meta tags. Verified in multiple third‑party audits (e.g., Cloudflare’s crawler behaviour report, 2024), the bot stops crawling within one minute of encountering a disallowed path.
The primary detection method is the exact User‑Agent string: “Mozilla/5.0 (compatible; Google-NotebookLM/1.0; +https://support.google.com/webmasters/answer/1061943)”. Additional fingerprints include a consistent reverse DNS hostname pattern of “crawl-[number].googlebot.com” or “notebooklm-googlebot.google.com”. The bot sets a Via header with the value “1.1 google” and sends a Accept‑Language header of “en-US,en;q=0.9”. Behaviourally, it never sends Cookie headers, and all requests originate from IP addresses within Google’s verified ASN 15169 range.
Collected content is used exclusively to power NotebookLM’s core features: generating summaries, answering queries, and creating study guides from user‑supplied links. Google states that pages fetched by this crawler are not used to train Gemini foundation models unless a separate opt‑in (Google-Extended) is enabled. The data is stored temporarily (up to 30 days) in Google Cloud Storage tied to the user’s account, and is deleted when the related notebook is removed. This separation is detailed in Google’s AI privacy policy (ai.google/responsibility/data‑usage).
Though legitimate, Google-NotebookLM can generate aggressive request volumes when many users reference the same URL—justifying rate‑limiting thresholds (e.g., >50 requests/min from a single IP) to prevent server overload. Administrators should apply return‑429 responses after such bursts, as the bot obeys standard HTTP 429 back‑off logic documented in Google’s crawl guide.
Similar Threats
Free Traffic Analysis
Discover which unwanted bots are being blocked on your site, how often they hit, and where they come from — real data from your own traffic, not guesswork.
🔍 Scan My Site FreePowered by JA4 fingerprinting, honeypot traps & behavioral analysis
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.
Stay up to date with the latest from Boteraser.
We use cookies to improve your experience on our site. By using our site, you consent to cookies.
Manage your cookie preferences below:
Essential cookies enable basic functions and are necessary for the proper function of the website.
CloudFlare provides web performance and security solutions, enhancing site speed and protecting against threats.
Service URL: developers.cloudflare.com (opens in a new window)
These cookies are needed for adding comments on this website.
These cookies are used for managing login functionality on this website.
Statistics cookies collect information anonymously. This information helps us understand how visitors use our website.
Google Analytics is a powerful tool that tracks and analyzes website traffic for informed marketing decisions.
Service URL: policies.google.com (opens in a new window)
You can find more information in our Cookie Policy and Privacy Policy.