formulafinderbot
formulafinderbot is a web crawler operated by Formula Finder, a B2B lead-generation platform headquartered in San Francisco, California, that provides a Chrome browser extension for extracting email addresses and contact data from publicly accessible websites. The bot’s primary purpose is to collect publicly listed email addresses, phone numbers, and social media links from web pages to populate Formula Finder’s proprietary lead database, which is then sold to sales and marketing teams. According to the company’s official website (formulafinder.com) and its robots.txt documentation, the crawler is designed exclusively to harvest contact information from public sources and does not access authenticated or private content.
formulafinderbot sends HTTP GET requests at a moderate frequency, typically one request every 2–3 seconds per domain, though this can increase to up to 10 requests per minute during targeted crawls. The bot identifies itself with the User-Agent string “formulafinderbot/1.0” and requests HTML pages, JavaScript files, and PDFs when scanning for email patterns. It does not follow links recursively beyond a depth of 3 levels, and it respects meta robots tags with nofollow and noindex directives. IP ranges for formulafinderbot are drawn from Amazon Web Services (AWS) EC2 instances, primarily in us-east-1 and us-west-2 regions, as detailed in a GitHub issue (#42) on the Formula Finder repository where webmasters reported the bot’s IP addresses ranging from 54.237.x.x to 52.44.x.x. The bot sends a User-Agent header only and does not include custom HTTP headers such as X-Forwarded-For or Referer.
According to Formula Finder’s official robots.txt policy page (archived at formulafinder.com/robots) and multiple forum discussions on WebmasterWorld, formulafinderbot is designed to honor Disallow directives in robots.txt files. In practice, however, several webmasters have reported that the bot occasionally ignores Disallow rules for pages containing email patterns, an issue acknowledged by Formula Finder in a 2019 blog post as a bug that was later patched. The company maintains a public list of IP addresses that can be blocked via firewall if robots.txt is not sufficient, as documented in their GitHub repository (github.com/formulafinder/crawler-policy).
The definitive indicator is the User-Agent string “formulafinderbot/1.0”, which may appear as Mozilla/5.0 (compatible; formulafinderbot/1.0; +https://formulafinder.com/bot) in server logs. Behavioral fingerprints include a consistent 2–3 second delay between requests, a high percentage of requests for contact, about, and team pages, and the absence of JavaScript parsing (the bot does not execute scripts). A secondary identifier is the IP range’s reverse DNS lookup showing ec2-*.*.compute-1.amazonaws.com. Formula Finder also provides a verification endpoint at https://formulafinder.com/bot/verify that returns a JSON object confirming the bot’s identity when called with the User-Agent.
Collected data—primarily email addresses, phone numbers, and LinkedIn profile URLs—is aggregated into Formula Finder’s lead database and sold to subscribers through their Chrome extension and API. The company claims to remove duplicates daily and to refresh its dataset every 30 days. A privacy policy (effective May 2022) states that only publicly visible information is stored and that users can request deletion via a webform. Data is not used for AI training; it is strictly for B2B sales intelligence and outbound marketing campaigns, as confirmed by Formula Finder’s support documentation.
Websites should rate-limit formulafinderbot to a maximum of 10 requests per minute per IP to maintain service stability, as the bot’s moderate crawl rate can still overwhelm small servers when multiple instances run concurrently. The recommended threshold for 429 responses is 50 requests per minute across all IPs, after which a 24-hour block is justified based on the bot’s documented behavior and the company’s own guidelines for excessive crawling.
Similar Threats
🛡️
Boteraser automatically detects and blocks unwanted bots — protecting your site from scrapers, DDoS bursts, and credential stuffing attacks without slowing down real visitors.
✅ Start Free ProtectionSetup takes under a minute · Free trial available
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.
Stay up to date with the latest from Boteraser.
We use cookies to improve your experience on our site. By using our site, you consent to cookies.
Manage your cookie preferences below:
Essential cookies enable basic functions and are necessary for the proper function of the website.
CloudFlare provides web performance and security solutions, enhancing site speed and protecting against threats.
Service URL: developers.cloudflare.com (opens in a new window)
These cookies are needed for adding comments on this website.
These cookies are used for managing login functionality on this website.
Statistics cookies collect information anonymously. This information helps us understand how visitors use our website.
Google Analytics is a powerful tool that tracks and analyzes website traffic for informed marketing decisions.
Service URL: policies.google.com (opens in a new window)
You can find more information in our Cookie Policy and Privacy Policy.