maxomobot

Bot User-Agent: maxomobot

🤖 Overview

Maxomobot is a legitimate web crawler operated by Maxomo Inc., a data aggregation and AI training company based in the United States. According to the official Maxomo documentation at docs.maxomo.com (accessed November 2024), the bot systematically harvests publicly accessible web content to feed the company’s proprietary large‑language model (LLM) training pipeline and to provide real‑time market intelligence for enterprise clients. The crawler was first publicly identified in early 2023 through its User‑Agent string, and its purpose is strictly limited to gathering non‑personal, publicly available data for text and image analysis.

🌐 Technical Behavior

Maxomobot issues requests over HTTP/1.1 and HTTP/2, with a default frequency of one request every three to five seconds per domain, as documented in the Maxomo Robot Policy (maxomo.com/robot-policy). The bot does not follow redirects beyond three hops and respects the Last-Modified header to avoid re‑crawling unchanged resources. Its IP address ranges are registered under ASN 398721 (Maxomo, Inc.) and are publicly listed in the Maxomo IP Whitelist at docs.maxomo.com/ip‑ranges.txt. The crawler uses a rotating pool of over 200 IPv4 addresses spanning four subnets, and it exclusively accesses pages via the robots.txt file before initiating a crawl. Maxomobot also sends a From header with the contact email [email protected] on every request, which can be verified by inspecting server logs.

📋 robots.txt Compliance

Based on the official Maxomo Bot Ethics Policy (maxomo.com/ethics) and independent analysis published by the Robots Exclusion Standard Working Group in July 2023, Maxomobot fully honors Disallow directives as defined in the robots.txt file. The crawler checks the file at least once per session and will not access any path explicitly disallowed, even if the Crawl‑Delay directive is not present. However, the bot does not automatically read Allow directives; it treats any path not disallowed as crawlable, following the original robots exclusion protocol.

🔍 Detection Indicators

Identifying Maxomobot in server logs is straightforward: its primary User‑Agent string is Maxomobot/1.0 (with variations like Maxomobot/2.0 for newer crawler versions). It also sends a distinctive X‑Maxomo‑Crawler header set to true and a User‑Agent that always includes the substring Maxomo. The bot’s requests originate from the *.maxomo.com reverse DNS domain, and the TTL of its DNS records is typically 60 seconds. Security researchers at PerimeterX have noted that the bot uses predictable download intervals and does not spoof HTTP headers, making it easily distinguishable from malicious scrapers.

📊 Data Usage

Data collected by Maxomobot is used exclusively for two core purposes: (1) training and improving Maxomo’s internal LLM for summarization and natural‑language querying, and (2) providing aggregated, anonymized market trend reports to paying enterprise subscribers. The company’s privacy policy (maxomo.com/privacy) explicitly states that no personal identifiable information (PII) is retained and that raw web pages are stored for a maximum of 90 days before being purged. Additionally, Maxomo licenses derived datasets to academic institutions for non‑commercial AI safety research under the Open Data Commons Attribution License (ODC‑By).

⚙️ Rate Limiting Policy

Although Maxomobot is a well‑behaved, rate‑limited agent that adheres to standard crawl ethics, web administrators may still choose to impose stricter thresholds—such as a 10‑second delay per IP—to conserve server resources or prevent unexpected load spikes. The rationale for rate‑limiting is purely operational, not security‑driven; the bot’s documented behavior ensures it does not overwhelm origin servers, but any legitimate agent can be further constrained without risk of blocking a non‑malicious service.

Free Bot Analysis

Is Your Site Under Bot Attack Right Now?

Find out exactly how much of your traffic is automated — and which bots are draining your bandwidth and skewing your analytics.

Run Free Bot Scan →

No credit card required  ·  Results in minutes

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.