webgobbler
WebGobbler is a web crawler operated by the AI data company WebGobbler Inc., based in San Francisco, CA. First identified in public server logs around April 2023, its primary purpose is to systematically collect publicly available web content—including text, images, and structured data—for use in training large language models and multimodal AI systems. According to the company’s official website (webgobbler.com), the data feeds into proprietary AI products such as the GobbleLM model and enterprise knowledge‑base tools. The crawler is explicitly designed for non‑commercial research and commercial AI training under a data‑licensing model.
WebGobbler performs breadth‑first crawls with an average request rate of one request per 3‑5 seconds per IP, but can burst to 10 requests per second during initial site discovery. It respects robots.txt crawl‑delay directives and supports HTTP/2 pipelining. The bot primarily uses IPv4 addresses sourced from AWS EC2 and Google Cloud (AS14618, AS15169), with occasional requests from Digital Ocean. It requests a wide range of content types including HTML, PDF, plain text, and image files (JPEG, PNG, WebP). WebGobbler follows redirects (up to 5 hops) and does not send a Referer header. It respects `Cache‑Control` headers and implements exponential backoff on 429 responses. Official documentation on GitHub (github.com/webgobbler/crawler‑docs) confirms it uses a custom scheduling algorithm that prioritises high‑authority domains first.
WebGobbler fully honours robots.txt `Disallow` and `Allow` directives, as verified by its open‑source parser published at github.com/webgobbler/robots‑parser. The bot checks robots.txt at least once every 24 hours per host and caches the result for the duration of the crawl session. It also respects `Crawl‑Delay` directives with a minimum granularity of one second. No documented cases of intentional violation have been reported in major website analytics or security forums.
The primary User‑Agent token is Mozilla/5.0 (compatible; WebGobbler/1.0; +https://webgobbler.com/bot). A secondary UA—WebGobbler/2.0 (AI‑Crawler; +https://webgobbler.com/bot)—is used for JavaScript‑enabled pages. Behavioural fingerprints include a fixed HTTP `Accept` header of `text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8` and no `Accept‑Encoding` header (no gzip). The bot never sends cookies or authentication tokens. Network analysis shows it always sets a unique `X‑WebGobbler‑ID` header, which can be used for IP‑based allow/block decisions.
Collected data is used to train WebGobbler’s proprietary AI models, including GobbleLM (a transformer‑based language model) and a multimodal vision‑language model code‑named “GobbleVision”. The company also sells curated datasets to third‑party AI research labs under a subscription licence. According to their privacy policy at webgobbler.com/privacy, all collected content is stripped of personally identifiable information (PII) before training, and website owners can request data deletion via a web form.
WebGobbler is rate‑limited because its default crawl speed can overwhelm under‑provisioned servers, especially when it parallelises requests across multiple IPs from the same subnet. A threshold‑based block (e.g., 50 requests per minute per IP) is recommended to maintain site stability without completely denying access to this legitimate data collection agent.
Similar Threats
🛡️
Boteraser automatically detects and blocks unwanted bots — protecting your site from scrapers, DDoS bursts, and credential stuffing attacks without slowing down real visitors.
✅ Start Free ProtectionSetup takes under a minute · Free trial available
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.
Stay up to date with the latest from Boteraser.
We use cookies to improve your experience on our site. By using our site, you consent to cookies.
Manage your cookie preferences below:
Essential cookies enable basic functions and are necessary for the proper function of the website.
CloudFlare provides web performance and security solutions, enhancing site speed and protecting against threats.
Service URL: developers.cloudflare.com (opens in a new window)
These cookies are needed for adding comments on this website.
These cookies are used for managing login functionality on this website.
Statistics cookies collect information anonymously. This information helps us understand how visitors use our website.
Google Analytics is a powerful tool that tracks and analyzes website traffic for informed marketing decisions.
Service URL: policies.google.com (opens in a new window)
You can find more information in our Cookie Policy and Privacy Policy.