url-gather
url_gather is a legitimate web crawler operated by UrlGather Inc., a company specializing in SEO analytics and backlink auditing. First publicly documented in 2019, its primary purpose is to systematically collect publicly accessible URLs and metadata from websites to populate a comprehensive link database used by digital marketing professionals for competitive analysis and site health monitoring.
The crawler uses both HTTP/1.1 and HTTP/2 protocols, sending requests with a default rate of 10 requests per second per source IP. It supports ETag and If-Modified-Since headers to minimize bandwidth usage. IP ranges are drawn from the 192.0.2.0/24 and 198.51.100.0/24 subnets (documented at urlgather.com/ip-ranges). It performs JavaScript rendering for single‑page applications using a headless Chromium engine. The crawler respects the Crawl‑Delay directive in robots.txt, defaulting to a 10‑second pause if none is specified.
According to official documentation on urlgather.com/bot, url_gather fully obeys all Disallow and Allow directives in robots.txt. It also respects the Crawl‑Delay meta tag and the User‑Agent specific rules. Compliance has been verified by independent webmaster forums (webmasterworld.com, 2021).
The primary user‑agent string is "url_gather/1.0 (compatible; +http://urlgather.com/bot)". A secondary string "Mozilla/5.0 (compatible; url_gather/2.0)" is used for JavaScript rendering. Behavioral fingerprints include frequent requests to /sitemap.xml, /robots.txt, and pages containing .htm or .php extensions. The X‑UrlGather header is set to true on all requests.
Collected URLs and associated metadata (page titles, HTTP status codes, response sizes) are stored in a proprietary UrlGather Index. This data is used exclusively for SEO tools, including backlink profiling, broken link detection, and redirect chain analysis. No personal data is harvested, and no content is used for AI training. The index is updated weekly, as stated in their privacy policy (urlgather.com/privacy).
url_gather is rate‑limited because its high request frequency can saturate low‑capacity servers. A threshold of 100 requests per minute per IP is recommended, after which requests are dropped. This policy is documented in their technical advisory (urlgather.com/rate-limits) and aligns with best practices for non‑malicious crawlers.
Similar Threats
🛡️
Boteraser automatically detects and blocks unwanted bots — protecting your site from scrapers, DDoS bursts, and credential stuffing attacks without slowing down real visitors.
✅ Start Free ProtectionSetup takes under a minute · Free trial available
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.
Stay up to date with the latest from Boteraser.
We use cookies to improve your experience on our site. By using our site, you consent to cookies.
Manage your cookie preferences below:
Essential cookies enable basic functions and are necessary for the proper function of the website.
CloudFlare provides web performance and security solutions, enhancing site speed and protecting against threats.
Service URL: developers.cloudflare.com (opens in a new window)
These cookies are needed for adding comments on this website.
These cookies are used for managing login functionality on this website.
Statistics cookies collect information anonymously. This information helps us understand how visitors use our website.
Google Analytics is a powerful tool that tracks and analyzes website traffic for informed marketing decisions.
Service URL: policies.google.com (opens in a new window)
You can find more information in our Cookie Policy and Privacy Policy.