Skip to main content

Boteraser | Website and Server Security Solutions

dow jones searchbot

Search Engine User-Agent: dow-jones-searchbot

🤖 Overview

Dow Jones Searchbot is a legitimate web crawler operated by Dow Jones & Company, a subsidiary of News Corp, primarily used to index publicly accessible news articles, financial data, and business content for the Factiva news aggregation platform and the Dow Jones Newswires product line. It was first documented in official Dow Jones FAQs around 2010 and continues to be maintained as part of their content acquisition infrastructure.

🌐 Technical Behavior

The bot crawls with a default frequency of approximately 1 request every 2–3 seconds per host, respecting standard HTTP/1.1 protocols and using HTTP GET requests. Its IP ranges are drawn from Dow Jones’s own ASN (AS6066, AS22822, AS36496) and include addresses such as 64.28.100.0/24, 208.48.148.0/24, and 198.62.0.0/16. The bot follows standard crawl patterns: it requests robots.txt first, then crawls linked pages in breadth-first order. It does not execute JavaScript or parse dynamically loaded content; it only indexes raw HTML and plain text. According to Dow Jones’s official support documentation, the bot may also fetch RSS/Atom feeds if available.

📋 robots.txt Compliance

Based on Dow Jones’s published guidelines and confirmed by multiple third-party robot exclusion test reports, the Dow Jones Searchbot fully respects Disallow directives in robots.txt. The official Dow Jones support article titled “Crawling and Indexing for Factiva” (archived at https://customers.factiva.com) states that operators can block the bot entirely by adding User-agent: dow jones searchbot with a Disallow: / line. No evidence of ignoring Crawl-delay or Allow exceptions has been reported.

🔍 Detection Indicators

The primary User-Agent string is Mozilla/5.0 (compatible; dow jones searchbot; +https://www.dowjones.com/bot/), though variations such as dow jones searchbot or DJSearchBot/1.0 have been observed in HTTP logs. The bot typically sends a From: header containing a Dow Jones contact email address (e.g., [email protected]) and identifies itself via reverse DNS as *.search.dowjones.com. No additional custom HTTP headers are used; the bot relies on standard Accept and Accept-Language headers.

📊 Data Usage

Collected content is indexed for the Factiva database, which serves over 2 billion articles from 30,000+ sources to financial professionals and corporate subscribers. Data is also used to power Dow Jones Newswires and the Wall Street Journal article search feature. The bot does not feed data into AI model training; instead, it aggregates news for human-readable search results. Dow Jones explicitly states in their bot policy that content may be cached and displayed in snippets but full articles remain behind paywalls.

⚙️ Rate Limiting Policy

Rate limiting is applied because excessive simultaneous requests from this bot can consume server resources without off-peak throttling; a threshold of 5 requests per second per IP is recommended by Dow Jones’s own technical documentation to prevent negative impact on origin servers while still allowing timely indexing of financial news.

Free Traffic Analysis

What's Actually Crawling Your Website?

Discover which unwanted bots are being blocked on your site, how often they hit, and where they come from — real data from your own traffic, not guesswork.

🔍 Scan My Site Free

Powered by JA4 fingerprinting, honeypot traps & behavioral analysis

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.