diffbot-user
Diffbot-User is the user-initiated crawler User-Agent operated by Diffbot, a California-based web data extraction company founded in 2012. It fetches pages on behalf of a Diffbot customer or API caller rather than as part of Diffbot’s autonomous Knowledge Graph crawl, feeding extracted content into Diffbot’s structured-data APIs and enrichment products. The bot is distinct from generic scrapers and is subject to Diffbot’s acceptable-use policies.
Requests appear as single-page or limited-batch fetches triggered by API calls to Diffbot's Article, Product, Image, or Discussion endpoints. Unlike the autonomous Diffbot crawler, Diffbot-User does not build a broad crawl frontier; it retrieves URLs explicitly submitted by users. Traffic originates from Diffbot’s cloud infrastructure, historically hosted on Amazon Web Services, and reverse DNS commonly resolves to diffbot.com domains. Diffbot has published guidance for verifying its crawler via reverse DNS and states that it does not publish a single static IP range because its infrastructure scales dynamically. It supports HTTP and HTTPS, follows redirects, and can render JavaScript for dynamic pages depending on the API configuration. The bot does not generally obey crawl-delay directives because each request is user-directed, but Diffbot applies internal concurrency limits to avoid overwhelming target sites.
Diffbot’s official bot documentation distinguishes the autonomous Diffbot crawler, which honors robots.txt Disallow rules, from Diffbot-User, which is documented as a user-initiated agent that may not honor robots.txt because a customer explicitly requested the URL. Site owners wishing to block the autonomous crawler can use User-agent: Diffbot, while blocking Diffbot-User may also affect legitimate API-driven requests.
Known User-Agent strings include Diffbot-User and variants containing Diffbot-User/1.0 or Diffbot/0.1 with the URL +http://www.diffbot.com. Identifying headers may include standard HTTP headers from AWS-hosted clients, and requests often target article, product, or listing pages with query parameters supplied by the API caller.
Extracted data is parsed by Diffbot’s machine-learning models into structured JSON, feeding the Diffbot Knowledge Graph, customer data pipelines, search and recommendation systems, and analytics workflows. The company states that it processes publicly accessible web content and provides APIs for entity extraction, article text, product data, and image analysis.
Rate limiting is appropriate because Diffbot-User can generate bursts when many API customers submit URLs simultaneously. Threshold-based blocking preserves access for legitimate low-volume users while preventing accidental site overload, consistent with Diffbot’s published crawler documentation at https://www.diffbot.com/bot/.
Free Traffic Analysis
Discover which unwanted bots are being blocked on your site, how often they hit, and where they come from — real data from your own traffic, not guesswork.
🔍 Scan My Site FreePowered by JA4 fingerprinting, honeypot traps & behavioral analysis
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.
Stay up to date with the latest from Boteraser.
We use cookies to improve your experience on our site. By using our site, you consent to cookies.
Manage your cookie preferences below:
Essential cookies enable basic functions and are necessary for the proper function of the website.
CloudFlare provides web performance and security solutions, enhancing site speed and protecting against threats.
Service URL: developers.cloudflare.com (opens in a new window)
These cookies are needed for adding comments on this website.
These cookies are used for managing login functionality on this website.
Statistics cookies collect information anonymously. This information helps us understand how visitors use our website.
Google Analytics is a powerful tool that tracks and analyzes website traffic for informed marketing decisions.
Service URL: policies.google.com (opens in a new window)
You can find more information in our Cookie Policy and Privacy Policy.