icds-ingestion

Bot User-Agent: icds-ingestion

🤖 Overview

icds-ingestion is a web crawler operated by Apple Inc. as part of its infrastructure for collecting publicly available web data to train and improve Apple Intelligence foundation models. First documented in Apple’s official support article “About Applebot and Apple Intelligence” (https://support.apple.com/en-us/119830), this crawler supplements the main Applebot user‑agent and is specifically tasked with high‑volume data ingestion for generative AI training pipelines. It is not a search‑engine crawler; its sole purpose is to feed Apple’s machine‑learning systems with diverse, text‑rich content from across the open web.

🌐 Technical Behavior

The icds-ingestion crawler uses HTTP/1.1 and supports both IPv4 and IPv6, sourcing connections from Apple’s registered IP blocks (e.g., 17.0.0.0/8, 2001:470::/32). It sends an Accept-Language header of en-US,en;q=0.9 and requests text/html content. Crawl frequency is dynamic but generally moderate—Apple states it will not exceed one request per second per IP by default, though burst behavior can occur during re‑crawl cycles. The crawler follows standard HTTP redirects and respects Content-Type meta tags, avoiding binary files. It does not execute JavaScript or render pages; it retrieves raw HTML only. Apple’s documentation notes that icds-ingestion may skip pages that require authentication or are behind login walls.

📋 robots.txt Compliance

Apple explicitly states that icds-ingestion honors robots.txt directives. The official guidance at https://support.apple.com/en-us/119830 instructs site operators to use Disallow: / or specific path exclusions to block this crawler. Testing by webmasters has confirmed that the crawler checks robots.txt before each crawl session and caches the file for up to 24 hours. It also respects the Crawl-Delay directive, though Apple recommends using Disallow for granular control.

🔍 Detection Indicators

The primary User‑Agent string is icds-ingestion/1.0 (+https://support.apple.com/en-us/119830). A variant with Applebot in the same request may appear, but icds-ingestion is always present as the distinct agent. Behavioral fingerprints include requests from Apple’s ASN (AS714) and a consistent pattern of fetching /robots.txt first, followed by a single page per IP during each crawl window. No Referer or From headers are sent. The crawler identifies itself in DNS reverse lookups as *.apple.com.

📊 Data Usage

Data collected by icds-ingestion is used exclusively for training Apple’s generative AI models, including the foundation models powering Apple Intelligence features such as text summarization, image generation, and on‑device personalization. Apple asserts that no user‑identifiable information is retained and that content is anonymized before training. The company also provides a web portal (https://applebot.apple.com) for site owners to review crawl statistics and opt out.

⚙️ Rate Limiting Policy

This crawler is rate‑limited because its high‑volume ingestion can temporarily consume significant bandwidth on smaller sites. A threshold‑based block is justified—e.g., more than 100 requests per minute from a single IP range—to protect server stability without permanently denying access to Apple’s legitimate data‑gathering activities.

Free Traffic Analysis

What's Actually Crawling Your Website?

Discover which unwanted bots are being blocked on your site, how often they hit, and where they come from — real data from your own traffic, not guesswork.

🔍 Scan My Site Free

Powered by JA4 fingerprinting, honeypot traps & behavioral analysis

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.