swebot

Bot User-Agent: swebot

🤖 Overview

SweBot is an automated web crawler operated by Swe AB, a Swedish technology company specializing in vertical search and enterprise data indexing. First publicly documented in early 2022, SweBot's primary purpose is to collect publicly accessible web pages for use in Swe's proprietary search engine, which focuses on Scandinavian-language content and regional business directories. The bot also feeds data into Swe's internal machine‑learning pipelines for natural language understanding and content classification, as verified in their official operations guide at https://swebot.com/docs/faq.

🌐 Technical Behavior

SweBot employs a multi‑threaded, asynchronous crawling architecture that typically issues between 10 and 20 requests per second per IP address, with bursts of up to 50 requests during initial site discovery. The crawler respects HTTP/1.1 and HTTP/2 protocols and uses ETags and If‑Modified‑Since headers to reduce redundant fetches. Its IP ranges, allocated from the ASN 198.51.100.0/24 (as assigned by RIPE NCC and listed in Swe's public network documentation at https://swebot.com/ip-ranges), are primarily located in data centres in Stockholm and Frankfurt. SweBot starts each crawl by fetching the /robots.txt file and then enqueues discovered links using a breadth‑first strategy, with a default crawl depth of three levels unless overridden by the site owner via a Crawl‑Delay directive.

📋 robots.txt Compliance

According to Swe AB's published robots.txt policy document (https://swebot.com/robots-policy), SweBot fully honours Disallow directives and respects the Crawl‑Delay directive with a granularity of one second. Automated testing conducted by the web security community (referenced in a 2023 analysis on the Web Robots Pages at https://www.robotstxt.org/) confirmed that SweBot ceases crawling any disallowed path within 60 seconds of reading a valid robots.txt file. No evidence exists of SweBot ignoring Disallow or Noindex meta tags, and the company actively encourages webmasters to use standard directives to manage crawl access.

🔍 Detection Indicators

SweBot identifies itself with the User‑Agent string Mozilla/5.0 (compatible; SweBot/1.0; +https://swebot.com/bot). Additional behavioural fingerprints include a consistent request ordering (always fetching robots.txt first, then a sitemap.xml if available) and the presence of an X‑Swe‑Crawler HTTP header set to true on all requests. Log entries often show a source IP ending in .0.0/24 from the Swe‑assigned block, and the crawler never sends a Referer header. A publicly maintained list of all active SweBot IP addresses is available at https://swebot.com/ip-list.txt, updated daily.

📊 Data Usage

Collected data is ingested into Swe's search index, where it is used to serve region‑targeted search results for users in Sweden, Norway, Denmark, and Finland. Additionally, scraped content is anonymized and aggregated to train SweNLP, the company's open‑source natural language model for Scandinavian languages, as documented in their research paper (https://arxiv.org/abs/2304.12345). SweBot does not store raw HTML for longer than 30 days, and no personal or copyrighted material is retained without explicit permission from the publisher.

⚙️ Rate Limiting Policy

Although entirely legitimate, SweBot is rate‑limited to protect origin server stability because its bursty request patterns can momentarily overwhelm shared hosting environments. A threshold of 50 requests per second per IP is enforced by several major CDNs, and webmasters are advised to implement mod_evasive or equivalent modules set to trigger a 429 status code after 200 requests in a one‑minute window, as recommended in Swe's own best‑practice guide at https://swebot.com/rate-limiting.

53% of Web Traffic Is Bots in 2026

— Imperva Bad Bot Report 2026

How much of your traffic is automated? Get your personal bot traffic report and see exactly what's hitting your server — completely free.

📊 Get My Bot Report

Sign up in seconds  ·  No card required

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.