Skip to main content

Boteraser | Website and Server Security Solutions

Diffbot-User

Bot User-Agent: diffbot-user

🤖 Overview

Diffbot-User is the user-initiated crawler User-Agent operated by Diffbot, a California-based web data extraction company founded in 2012. It fetches pages on behalf of a Diffbot customer or API caller rather than as part of Diffbot’s autonomous Knowledge Graph crawl, feeding extracted content into Diffbot’s structured-data APIs and enrichment products. The bot is distinct from generic scrapers and is subject to Diffbot’s acceptable-use policies.

🌐 Technical Behavior

Requests appear as single-page or limited-batch fetches triggered by API calls to Diffbot's Article, Product, Image, or Discussion endpoints. Unlike the autonomous Diffbot crawler, Diffbot-User does not build a broad crawl frontier; it retrieves URLs explicitly submitted by users. Traffic originates from Diffbot’s cloud infrastructure, historically hosted on Amazon Web Services, and reverse DNS commonly resolves to diffbot.com domains. Diffbot has published guidance for verifying its crawler via reverse DNS and states that it does not publish a single static IP range because its infrastructure scales dynamically. It supports HTTP and HTTPS, follows redirects, and can render JavaScript for dynamic pages depending on the API configuration. The bot does not generally obey crawl-delay directives because each request is user-directed, but Diffbot applies internal concurrency limits to avoid overwhelming target sites.

📋 robots.txt Compliance

Diffbot’s official bot documentation distinguishes the autonomous Diffbot crawler, which honors robots.txt Disallow rules, from Diffbot-User, which is documented as a user-initiated agent that may not honor robots.txt because a customer explicitly requested the URL. Site owners wishing to block the autonomous crawler can use User-agent: Diffbot, while blocking Diffbot-User may also affect legitimate API-driven requests.

🔍 Detection Indicators

Known User-Agent strings include Diffbot-User and variants containing Diffbot-User/1.0 or Diffbot/0.1 with the URL +http://www.diffbot.com. Identifying headers may include standard HTTP headers from AWS-hosted clients, and requests often target article, product, or listing pages with query parameters supplied by the API caller.

📊 Data Usage

Extracted data is parsed by Diffbot’s machine-learning models into structured JSON, feeding the Diffbot Knowledge Graph, customer data pipelines, search and recommendation systems, and analytics workflows. The company states that it processes publicly accessible web content and provides APIs for entity extraction, article text, product data, and image analysis.

⚙️ Rate Limiting Policy

Rate limiting is appropriate because Diffbot-User can generate bursts when many API customers submit URLs simultaneously. Threshold-based blocking preserves access for legitimate low-volume users while preventing accidental site overload, consistent with Diffbot’s published crawler documentation at https://www.diffbot.com/bot/.

Free Traffic Analysis

What's Actually Crawling Your Website?

Discover which unwanted bots are being blocked on your site, how often they hit, and where they come from — real data from your own traffic, not guesswork.

🔍 Scan My Site Free

Powered by JA4 fingerprinting, honeypot traps & behavioral analysis

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.