Skip to main content

Boteraser | Website and Server Security Solutions

webgains-bot

Bot User-Agent: webgains-bot

🤖 Overview

webgains-bot is a legitimate crawler operated by Webgains Ltd., a UK-based affiliate marketing network founded in 2004 and headquartered in Nottingham. The bot’s primary purpose is to scan participating publisher websites to verify the presence and correct placement of Webgains affiliate links, promotional banners, and tracking codes, ensuring compliance with the network’s terms of service. It also indexes product pages and campaign content to feed Webgains’ proprietary analytics platform, which advertisers and publishers use to monitor click-through rates and sales conversions. According to Webgains’ official documentation (https://www.webgains.com/public/en/terms/), the bot is part of their automated quality assurance system.

🌐 Technical Behavior

The bot performs periodic HTTP GET requests, typically starting from a seed list of known affiliate partner URLs. It follows internal links to a depth of 2–3 pages per session, focusing on paths containing keywords like “product”, “coupon”, or “affiliate”. Request frequency is moderate, with an average of 1–3 requests per second per domain, though burst traffic may occur during initial site audits. IP ranges are drawn from a static set of Amazon Web Services (AWS) and Linode blocks, as confirmed by reverse DNS lookups (e.g., ec2-*-*-*-*.eu-west-1.compute.amazonaws.com). The bot uses HTTP/1.1 with default headers and does not support gzip compression in all requests. It intentionally avoids crawling admin paths, login pages, and /wp-admin directories by default, as documented in the Webgains publisher integration guide.

📋 robots.txt Compliance

Webgains officially states that webgains-bot respects robots.txt Disallow directives, but field tests by site operators (reported on forums like WebmasterWorld) show inconsistent behaviour when custom rules are applied to subdirectories. For example, a Disallow: /secret/ rule is honoured within 24–48 hours, while rules targeting specific file extensions (e.g., *.pdf) may be ignored. The bot’s official documentation recommends using the user-agent line User-agent: webgains-bot in robots.txt to ensure compliance.

🔍 Detection Indicators

The primary User‑Agent string is webgains-bot (compatible; Webgains; +https://www.webgains.com/bot.html). A secondary variant Webgains/1.0 may appear in legacy logs. The bot sends a unique header X-Webgains-Bot: true on all requests. Behavioral fingerprints include a request pattern that always appends a parameter ?wgref= with a numeric publisher ID. Missing Referer headers and the absence of JavaScript execution are also characteristic.

📊 Data Usage

Collected data—including page URLs, link anchor text, product prices, and affiliate code presence—is used exclusively for internal quality assurance and affiliate commission verification. Webgains does not sell or repurpose the data for AI training or third-party analytics. The information feeds their tracking dashboard, enabling real-time error alerts for broken links or misconfigured tags. According to Webgains’ privacy policy (https://www.webgains.com/public/en/privacy/), the bot does not store personal information or session cookies.

⚙️ Rate Limiting Policy

Rate limiting is applied to webgains-bot because its periodic scans, while non‑malicious, can spike during global campaign launches (e.g., Black Friday), overwhelming small sites. A sensible threshold—such as 20 requests per 60 seconds per IP—preserves site performance while allowing the bot to complete its essential compliance checks within a reasonable timeframe.

Free Bot Analysis

Is Your Site Under Bot Attack Right Now?

Find out exactly how much of your traffic is automated — and which bots are draining your bandwidth and skewing your analytics.

Run Free Bot Scan →

No credit card required  ·  Results in minutes

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.