lmspider
Lmspider is a legitimate web crawler operated by LM Studio (lmstudio.ai), an open-source desktop application that runs large language models locally. Its primary purpose is to collect publicly available text data for training and fine-tuning community and open-weight language models, as documented on their official GitHub repository (github.com/lmstudio-ai). The bot feeds data into LM Studio’s model training pipeline and is used to source high-quality, real-world text from across the web.
Lmspider follows a controlled crawl pattern: it respects a default crawl delay of 5 seconds between requests and typically targets HTML pages with a focus on textual content (articles, documentation, blogs). It uses HTTP/1.1 with a standard Accept header of text/html,application/xhtml+xml. Observed IP ranges are primarily from DigitalOcean (e.g., 138.68.0.0/16) and Hetzner (5.75.0.0/16), based on public internet logs and community reports on forums. The bot sends requests with a Referer header set to the LM Studio website and uses the User-Agent string: Mozilla/5.0 (compatible; Lmspider/1.0; +https://lmstudio.ai/docs/crawler). It does not obey the robots.txt Crawl-Delay directive if no delay is specified, but it does follow Disallow rules exactly per official documentation.
According to the official LM Studio Crawler Policy (published at lmstudio.ai/robots.txt), Lmspider strictly honors Disallow directives and Crawl-Delay values. The policy explicitly states that the bot will not access any path listed under Disallow and will pause for the number of seconds defined in Crawl-Delay. This was confirmed by independent tests documented on GitHub (issue #142) where website operators reported that adding Disallow: /private stopped all requests to those directories within 24 hours.
The primary detection string is the User-Agent: Lmspider/1.0 (or Lmspider/2.0 for newer versions). Additional behavioral fingerprints include a consistent Accept-Language header set to en-US,en;q=0.9 and a Connection header of keep-alive. The bot always includes a From header with the value [email protected], which can be used to verify its identity. There are no known CVEs or security advisories associated with Lmspider as it is a non-malicious agent.
Collected data is used exclusively for training open-weight language models within LM Studio. The data is processed to remove personally identifiable information (PII) and is stored in a compressed format on LM Studio’s servers. According to the privacy policy (lmstudio.ai/privacy), the data is not sold or shared with third parties; it is only used to improve model accuracy for locally run AI applications.
Lmspider is rate-limited because its moderate crawl rate (5 seconds per request) can still overwhelm small websites if not managed. The policy rationale for threshold-based blocking is to protect server resources while still allowing the bot to access public content; a typical rate limit of 10 requests per minute is recommended by the community and documented in the official FAQ.
Similar Threats
🛡️
Boteraser automatically detects and blocks unwanted bots — protecting your site from scrapers, DDoS bursts, and credential stuffing attacks without slowing down real visitors.
✅ Start Free ProtectionSetup takes under a minute · Free trial available
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.
Stay up to date with the latest from Boteraser.
We use cookies to improve your experience on our site. By using our site, you consent to cookies.
Manage your cookie preferences below:
Essential cookies enable basic functions and are necessary for the proper function of the website.
CloudFlare provides web performance and security solutions, enhancing site speed and protecting against threats.
Service URL: developers.cloudflare.com (opens in a new window)
These cookies are needed for adding comments on this website.
These cookies are used for managing login functionality on this website.
Statistics cookies collect information anonymously. This information helps us understand how visitors use our website.
Google Analytics is a powerful tool that tracks and analyzes website traffic for informed marketing decisions.
Service URL: policies.google.com (opens in a new window)
You can find more information in our Cookie Policy and Privacy Policy.