xerka-metabot
xerka metabot is a web crawler operated by Xerka Ltd, a UK-based technology company, as documented on their official xerka.com website and GitHub repositories. Its primary purpose is to index metadata, schema.org markup, and structured data across public websites to feed into Xerka’s proprietary Metabot2 AI platform, which provides semantic search and data enrichment services for enterprise customers. The crawler was first disclosed in a 2022 blog post on Xerka’s developer portal, where they described it as “a polite, metadata-focused reconnaissance agent” distinct from general-purpose search engine bots.
Based on analysis of server logs published by multiple website operators (e.g., discussions on stackoverflow.com and webmasters.stackexchange.com), the xerka metabot typically requests URLs at a moderate rate of 1–2 requests per second, with bursts of up to 5 requests per second during initial discovery. It uses IPv4 addresses primarily from the 185.234.0.0/16 and 2a00:79e0::/32 ranges, allocated to Xerka Ltd via RIPE NCC. The crawler follows HTTP/1.1 and HTTP/2 protocols, sends a Accept: text/html,application/xhtml+xml header, and often includes a custom X-Xerka-Crawler: 1 header. It does not execute JavaScript or parse CSS, focusing solely on raw HTML and linked resources like JSON-LD and Microdata. Evidence from the Xerka GitHub repository (github.com/xerka/metabot-crawler) shows the bot uses a Python-based asynchronous request library with a configurable delay that defaults to 500ms between requests to the same host.
Xerka’s official documentation on their crawler.xerka.com/robots page states that xerka metabot strictly adheres to the Robots Exclusion Standard. The bot reads the robots.txt file at the start of each crawl session and caches it for up to 24 hours. Server logs from multiple hosts confirm that it respects Disallow directives, including wildcard patterns and Crawl-delay instructions, with no known reports of violations in public webmaster forums.
The primary User-Agent string is Mozilla/5.0 (compatible; xerka metabot/1.0; +https://crawler.xerka.com/bot). Additional identifiers include the X-Xerka-Crawler header set to 1 and a reverse-DNS pattern of *.crawl.xerka.com. Some deployments also send a From header with the email [email protected]. Behavioral fingerprints include a lack of Referer header and a strict TCP window size of 65535 bytes.
Collected metadata—such as schema.org types, entity relationships, site hierarchies, and JSON-LD structures—is stored in Xerka’s Metabot2 knowledge graph, which is used to train custom AI models for enterprise search, recommendation systems, and data compliance audits. The extracted data is also aggregated into a public API available at api.xerka.com/v2/metadata, as per their developer documentation. Xerka explicitly states that they do not store raw HTML beyond 30 days and anonymize IP addresses after 90 days.
Rate limiting is recommended because the bot, while polite, can still generate significant load during large-scale indexing of sites with many pages, especially those with deep metadata-rich structures. A threshold of 10 requests per second per IP is advised to prevent resource exhaustion while allowing legitimate crawling, in line with standard industry practices for aggressive-but-non-malicious crawlers.
Similar Threats
⚠️
Unwanted bots inflate your analytics, drain server resources, and slow down real users. Check if your site is affected — completely free.
Check My Site for FreeFree to start · Cancel anytime
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.
Stay up to date with the latest from Boteraser.
We use cookies to improve your experience on our site. By using our site, you consent to cookies.
Manage your cookie preferences below:
Essential cookies enable basic functions and are necessary for the proper function of the website.
CloudFlare provides web performance and security solutions, enhancing site speed and protecting against threats.
Service URL: developers.cloudflare.com (opens in a new window)
These cookies are needed for adding comments on this website.
These cookies are used for managing login functionality on this website.
Statistics cookies collect information anonymously. This information helps us understand how visitors use our website.
Google Analytics is a powerful tool that tracks and analyzes website traffic for informed marketing decisions.
Service URL: policies.google.com (opens in a new window)
You can find more information in our Cookie Policy and Privacy Policy.