image-fetcher
Bot User-Agent:image-fetcher
🤖 Overview
image-fetcher is a legitimate web crawler operated by DuckDuckGo, used to fetch and cache images for display in DuckDuckGo image search results. First documented in 2019, its sole purpose is to retrieve image data (JPEG, PNG, GIF, WebP) from public web pages to build the search engine’s image index.
🌐 Technical Behavior
The image-fetcher bot sends HTTP GET requests with a variable crawl rate, typically 1–5 requests per second per IP, but can burst up to 20 requests in short intervals during initial index updates. It uses IPv4 ranges published in DuckDuckGo’s official IP list (e.g., 50.18.x.x and 54.193.x.x, as confirmed in DuckDuckGo’s crawler documentation at duckduckgo.com/duckduckgo-help-pages/results/sources/). The bot fetches only images, not HTML pages, and does not execute JavaScript or follow links. It requests images with Accept: image/* and ignores Content-Type headers other than image types. The User-Agent string includes the bot’s version and a contact URL for feedback (see official DuckDuckGo crawler page).
📋 robots.txt Compliance
According to DuckDuckGo’s official guide, image-fetcher fully respects robots.txt Disallow directives, including the User-agent: DuckDuckGo-Image-Fetcher line. It also obeys Crawl-Delay instructions if specified, though DuckDuckGo recommends site owners use Disallow for images they don’t want indexed.
🔍 Detection Indicators
The User-Agent string is DuckDuckGo-Image-Fetcher/1.0 (+https://duckduckgo.com/duckduckgo-help-pages/results/sources/). The bot also sends the header From: [email protected] (a generic DuckDuckGo contact) and Accept-Encoding: gzip. Reverse DNS lookups for its IPs resolve to *.duckduckgo.com.
📊 Data Usage
Collected images are stored temporarily in DuckDuckGo’s image cache and used solely to generate thumbnail previews in search results. The images are not used for AI training or any other product; they are discarded after the cache expires (typically 7–14 days).
⚙️ Rate Limiting Policy
image-fetcher is rate-limited because high volumes of requests during initial index crawls can strain server resources. A threshold of 20 requests per 10 seconds per IP is recommended to prevent overloading while still allowing the bot to update the image index.
Similar Threats
🛡️
Stop Bots. Save Bandwidth. Protect Revenue.
Boteraser automatically detects and blocks unwanted bots — protecting your site from scrapers, DDoS bursts, and credential stuffing attacks without slowing down real visitors.
✅ Start Free ProtectionSetup takes under a minute · Free trial available
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.