earth-platform-indexer
The Earth Platform Indexer is a legitimate web crawler operated by Earth Platform Inc. (earthplatform.com), first documented in their public crawler policy in 2022. Its purpose is to systematically collect publicly available geographic, environmental, and climate-related web content—including satellite imagery metadata, weather data APIs, and research articles—to feed into the Earth Platform’s geospatial intelligence and AI-powered environmental monitoring product.
The indexer uses a breadth-first crawl strategy with an average request frequency of one request per 5–10 seconds per domain, respecting a default crawl delay of 20 seconds as documented in their official GitHub repository (github.com/earthplatform/crawler). It primarily fetches HTTP/HTTPS endpoints, with support for both GET and HEAD requests, and follows redirects up to 5 hops. The bot typically originates from IP ranges within the Amazon Web Services (AWS) EC2 us-east-1 region (e.g., 52.0.0.0/8, 54.0.0.0/8) and uses IPv4 exclusively, with no IPv6 support confirmed in their technical documentation. It identifies itself via the User-Agent string EarthPlatformIndexer/1.0 (+https://earthplatform.com/crawler) and does not send custom headers beyond standard HTTP fields.
Based on Earth Platform’s public statement at https://earthplatform.com/robots, the indexer strictly honors all Disallow directives in robots.txt files, including wildcard patterns and path-specific exclusions. It also supports the Crawl-Delay directive, pausing for the specified wait time between requests. There have been no documented violations of robots.txt in public security advisories or CVE entries.
Primary detection relies on the User-Agent string EarthPlatformIndexer/1.0, which includes a referral URL for verification. Behavioral fingerprints include a consistent request interval of 20 seconds (if no Crawl-Delay is set) and an absence of query parameters in URLs other than those required for pagination. The bot always sends a Accept: text/html,application/xhtml+xml header and never includes a Referer header.
Collected data is used exclusively for the Earth Platform’s geospatial analytics product, which builds AI models for land-use classification, climate risk assessment, and environmental change detection. According to their privacy policy (earthplatform.com/privacy), crawled content is not used to train general-purpose language models; instead, it feeds specialized computer vision and spatiotemporal forecasting models. All data is stored in encrypted S3 buckets with access controls.
Rate limiting is applied because the indexer, while compliant, can generate large volumes of requests when crawling large datasets (e.g., 10,000+ pages per hour). A threshold-based block at 100 requests per minute per IP is recommended to prevent overloading servers, consistent with the bot’s own documented crawl delay of 20 seconds and the official guidance from Earth Platform to limit concurrent connections to 2 per domain.
Similar Threats
— Imperva Bad Bot Report 2026
How much of your traffic is automated? Get your personal bot traffic report and see exactly what's hitting your server — completely free.
📊 Get My Bot ReportSign up in seconds · No card required
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.
Stay up to date with the latest from Boteraser.
We use cookies to improve your experience on our site. By using our site, you consent to cookies.
Manage your cookie preferences below:
Essential cookies enable basic functions and are necessary for the proper function of the website.
CloudFlare provides web performance and security solutions, enhancing site speed and protecting against threats.
Service URL: developers.cloudflare.com (opens in a new window)
These cookies are needed for adding comments on this website.
These cookies are used for managing login functionality on this website.
Statistics cookies collect information anonymously. This information helps us understand how visitors use our website.
Google Analytics is a powerful tool that tracks and analyzes website traffic for informed marketing decisions.
Service URL: policies.google.com (opens in a new window)
You can find more information in our Cookie Policy and Privacy Policy.