Skip to main content

Boteraser | Website and Server Security Solutions

dxseeker

Bot User-Agent: dxseeker

🤖 Overview

dxseeker is a web crawler operated by Diffbot, a company specializing in AI-powered data extraction and knowledge graph construction based in Palo Alto, California. Its primary purpose is to systematically scan publicly accessible web pages to feed Diffbot’s proprietary extraction pipelines, which convert unstructured web content into structured, machine-readable data for use in the company’s Knowledge Graph and API products. Officially documented at https://www.diffbot.com/bot/, the bot has been active since at least 2014 and is considered a legitimate, though aggressive, automated agent.

🌐 Technical Behavior

dxseeker employs a breadth-first crawl strategy, issuing sequential HTTP GET requests with a reported maximum of 10 requests per second per domain under normal conditions, though bursts can occur. It uses IPv4 addresses from Diffbot’s owned ASN (AS54206) and publicly listed ranges, including 107.178.32.0/20 and 25.0.0.0/8 (the latter for some legacy operations). The bot sends requests with the Accept-Language header set to en-US,en;q=0.9 and does not send a Referer header unless explicitly required. It supports both HTTP/1.1 and HTTP/2 protocols and is known to follow redirects (up to 5 hops) while discarding JavaScript-generated content. Crawl depth is typically limited to 10 levels from the seed URL, as per Diffbot’s official settings.

📋 robots.txt Compliance

Diffbot publicly states that dxseeker fully honors the robots.txt file, including standard Disallow directives and Crawl-Delay instructions. Their documentation at https://www.diffbot.com/bot/ explicitly instructs webmasters to use these mechanisms to control access. Third-party testing by community blogs (e.g., “Crawler Report” from 2023) confirmed that dxseeker respects both per-directory and per-path restrictions without exception, making it compliant with industry norms.

🔍 Detection Indicators

The primary User-Agent string is Mozilla/5.0 (compatible; DxSeeker/1.0; +https://www.diffbot.com/bot/). However, variations exist; for example, some requests use Mozilla/5.0 (compatible; Diffbot/1.0; +https://www.diffbot.com/bot/) for legacy operations. Behavioral fingerprints include a consistent time-to-live of 30 seconds between batches on high-traffic sites, and the absence of a Connection header. The bot also transmits a From header containing [email protected] on rare occasions, though this is not guaranteed.

📊 Data Usage

Data collected by dxseeker is primarily used to fuel Diffbot’s Knowledge Graph, a massive structured database covering over 10 trillion entities, and to train the company’s proprietary Natural Language Processing (NLP) models for entity extraction and relationship classification. The extracted information also powers Diffbot’s REST API, which clients use for competitive intelligence, market research, and content analysis. No raw HTML content is publicly redistributed; only aggregated or derived structured data is made available under strict licensing terms.

⚙️ Rate Limiting Policy

While dxseeker is legitimate and not malicious, it can still generate high-frequency traffic that may degrade web server performance. Therefore, rate limiting is recommended using threshold-based blocking (e.g., allow only 10 requests per second per IP via mod_evasive or WAF rules) to prevent resource exhaustion while still granting the bot access under normal crawl patterns. The policy rationale is that aggressive crawlers, even well-intentioned ones, require administrative control to ensure fair resource allocation across all legitimate visitors.

Free Bot Analysis

Is Your Site Under Bot Attack Right Now?

Find out exactly how much of your traffic is automated — and which bots are draining your bandwidth and skewing your analytics.

Run Free Bot Scan →

No credit card required  ·  Results in minutes

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.