atlocalbot

Bot User-Agent: atlocalbot

🤖 Overview

atlocalbot is a web crawler operated by AtLocal Ltd, a UK-based company founded in 2003 that powers local search and business directory services. According to the official AtLocal website and its bot policy page (http://www.atlocal.com/bot.html), the crawler’s primary purpose is to collect publicly available business information—such as address, phone number, opening hours, and reviews—for inclusion in AtLocal’s local search engine and directory products. It is a legitimate, non-malicious agent designed to aggregate local business data from multiple sources.

🌐 Technical Behavior

atlocalbot performs deep, systematic crawls of business directories, review sites, and general web pages that contain location-based content. It typically follows HTTP/1.1 with persistent connections and may request a high volume of pages per second (often 20–50 requests concurrently) from a single IP address. Official documentation from AtLocal indicates the bot uses IP ranges owned by AtLocal Ltd and hosted on UK-based servers, though some crawls may originate from cloud providers. It respects Cache-Control and ETag headers to avoid redundant downloads. Historically, the bot has been observed to ignore robots.txt disallow directives under certain conditions, but AtLocal states it obeys them after a short delay. The crawler does not crawl JavaScript-heavy content and relies on raw HTML.

📋 robots.txt Compliance

AtLocal’s official bot page (http://www.atlocal.com/bot.html) claims that atlocalbot “honours robots.txt directives” and recommends webmasters add a User-agent: atlocalbot line to disallow paths. However, anecdotal reports from webmasters and discussion forums (e.g., Stack Exchange’s Webmasters section, 2010–2023) note that the bot sometimes continues crawling after a disallow directive, albeit at a reduced rate. The company has acknowledged these issues and advises emailing support for persistent problems. In practice, the bot is documented to respect Disallow patterns after an initial burst of requests, making it moderately compliant.

🔍 Detection Indicators

The canonical User-Agent string for atlocalbot is: Mozilla/5.0 (compatible; atlocalbot/1.0; +http://www.atlocal.com/bot.html). Some variants include version numbers like 1.2 or 1.3. Bot traffic typically lacks common browser headers such as Accept-Language or Referer. The bot’s IP ranges are not officially published, but WHOIS lookups show origin ASNs belonging to AtLocal Ltd (AS395092). It does not use the X-Forwarded-For header. Traffic peaks during business hours in the UK (UTC+0/+1).

📊 Data Usage

Collected data is aggregated into AtLocal’s local search index and business directory, which is available via atlocal.com and third-party syndication partners. According to AtLocal’s privacy policy, the data is used to build a comprehensive, free directory of local businesses for end users and to power targeted local advertising. The crawler’s output feeds into machine learning models that deduplicate and categorize business listings, but not for general-purpose AI training. No personal data beyond publicly listed business information is collected.

⚙️ Rate Limiting Policy

atlocalbot is rate-limited because its high concurrency (up to 50 simultaneous connections) can overwhelm smaller websites, causing degraded performance for human visitors. The policy rationale for threshold-based blocking is to protect server resources while still allowing the bot to index legitimate business listings, as recommended in the official AtLocal guidelines. Webmasters should implement a rate limit of 10 requests per second per IP using server-level tools (e.g., Nginx limit_req) rather than outright blocking.

Free Traffic Analysis

What's Actually Crawling Your Website?

Discover which unwanted bots are being blocked on your site, how often they hit, and where they come from — real data from your own traffic, not guesswork.

🔍 Scan My Site Free

Powered by JA4 fingerprinting, honeypot traps & behavioral analysis

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.