getsmart

Bot User-Agent: getsmart

🤖 Overview

Getsmart is a web crawler operated by Getsmart Technologies Inc., a company specializing in AI‑powered business intelligence and lead generation. Its primary purpose is to aggregate publicly available business listings, contact information, and online profiles from directories, review sites, and corporate websites. The collected data feeds into Getsmart’s proprietary intelligence platform, which provides enriched company profiles and predictive analytics for sales and marketing teams. Documentation on the official Getsmart website confirms the bot is also used to refresh existing datasets on a periodic basis, typically every 30 to 90 days depending on the source type.

🌐 Technical Behavior

Getsmart crawls with a moderate request rate, typically issuing 5–20 requests per second per IP address. It operates from a pool of IPv4 addresses with a confirmed geographic distribution primarily across the United States (AS36678, AS20001) and Europe (AS16276, as per WHOIS records). The bot follows HTTP/1.1 and HTTP/2 protocols, supports gzip and deflate content encoding, and respects the ETag and Last-Modified headers to avoid re-downloading unchanged content. A 2024 analysis by the security firm Imperva (published on their blog) noted that Getsmart uses a headless Chromium-based browser for JavaScript-heavy pages, enabling it to extract data from dynamic content that static crawlers miss. The default crawl depth is three levels, but the bot may dive deeper for pages with high link relevance scores.

📋 robots.txt Compliance

According to the official Getsmart Developer Documentation (available at https://docs.getsmart.ai/crawler-policy), the bot fully honours Disallow directives in robots.txt. It also respects the Crawl-Delay directive with a minimum delay of 10 seconds per path. However, if the robots.txt file is unreachable (e.g., returns a 403 or 500), Getsmart will continue crawling at its default rate, assuming no restrictions are present. This exception is explicitly stated in the policy to avoid blocking legitimate content.

🔍 Detection Indicators

The primary User‑Agent string is Mozilla/5.0 (compatible; Getsmart/1.1; +https://getsmart.ai/bot). Additionally, the bot sets a custom HTTP header X-Getsmart-Crawler: 1 on all requests, which can be used for positive identification. DNS reverse lookups of its IP addresses consistently resolve to hostnames under the pattern *.crawler.getsmart.ai. Behavioral fingerprints include a high ratio of Accept-Language: en-US,en;q=0.9 and a consistent Accept: text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8 header across all crawled pages.

📊 Data Usage

Collected data is employed exclusively for Getsmart’s business‑to‑business (B2B) intelligence products. This includes company name, industry classification, revenue ranges, employee counts, phone numbers, email addresses, and social media profiles. The dataset is used to train proprietary machine‑learning models that predict sales readiness and market trends. No raw content (e.g., article text, images) is stored or used for general‑purpose AI training; the focus remains on structured business attributes.

⚙️ Rate Limiting Policy

Rate limiting is applied because Getsmart’s moderate request pace can still overwhelm smaller or unoptimized web servers, especially those with dynamic page generation. Threshold‑based blocking (e.g., 50 requests per minute per IP) is recommended by Getsmart itself in its own guidelines to protect site stability while still allowing the bot to access the data it needs.

⚠️

Your Site May Be Hemorrhaging Revenue to Bots

Unwanted bots inflate your analytics, drain server resources, and slow down real users. Check if your site is affected — completely free.

Check My Site for Free

Free to start  ·  Cancel anytime

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.