obot
oBot is a web crawler operated by Outbrain Inc., a global content recommendation platform. Its primary purpose is to index publicly available web content—such as article headlines, images, and excerpts—to populate the personalized widget recommendations displayed on thousands of publisher partner sites. The bot was first introduced around 2010 and has since evolved to support encrypted connections (HTTPS), HTTP/2, and dynamic content rendering via JavaScript.
oBot issues requests at a high volume, often hundreds per minute, from a large set of IP addresses allocated to Outbrain’s AWS infrastructure. Official Outbrain documentation lists IP ranges like 52.84.0.0/15 and 50.17.0.0/16, though these may change. The crawler uses the GET method, sends standard headers including Accept: text/html,application/xhtml+xml, and requests gzip encoding. It respects the Crawl-delay directive in robots.txt and also uses conditional If-Modified-Since headers to avoid re-fetching unchanged content. The bot typically checks for a robots.txt file before crawling any subdirectory.
According to Outbrain’s official robots.txt policy posted at https://www.outbrain.com/legal/robots/, the oBot crawler fully adheres to Disallow rules and will not access any URL path explicitly forbidden by the site’s robots.txt. It also respects the Crawl-delay parameter, pausing between requests as specified. This compliance is documented in webmaster forums and in Outbrain’s support pages, confirming the bot’s cooperative behavior.
The primary User-Agent string is oBot (case-insensitive), sometimes accompanied by a version suffix like oBot/1.0. Outbrain also uses the strings Outbrain Bot and Outbrain-ImageCrawler for image-specific crawling. Reverse DNS lookups on oBot IPs typically resolve to a domain ending in .outbrain.com or .awsdns.com. Behaviorally, the bot consistently requests one page per five seconds unless a different crawl delay is specified.
The data collected—page titles, meta descriptions, featured images, and text snippets—are used to build Outbrain’s content recommendation index. This index drives the machine learning algorithms that predict which articles a user is most likely to engage with. Outbrain’s privacy policy emphasizes that the bot does not intentionally collect personally identifiable information, focusing only on publicly accessible content to serve contextual recommendations on partner sites.
oBot is rate limited because its aggressive crawl pace can strain smaller servers. A typical threshold of 10 requests per second per source IP is sufficient to ensure the bot does not degrade site performance while still allowing Outbrain to fulfill its legitimate business need of discovering new content for its recommendation engine.
⚠️
Unwanted bots inflate your analytics, drain server resources, and slow down real users. Check if your site is affected — completely free.
Check My Site for FreeFree to start · Cancel anytime
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.
Stay up to date with the latest from Boteraser.
We use cookies to improve your experience on our site. By using our site, you consent to cookies.
Manage your cookie preferences below:
Essential cookies enable basic functions and are necessary for the proper function of the website.
CloudFlare provides web performance and security solutions, enhancing site speed and protecting against threats.
Service URL: developers.cloudflare.com (opens in a new window)
These cookies are needed for adding comments on this website.
These cookies are used for managing login functionality on this website.
Statistics cookies collect information anonymously. This information helps us understand how visitors use our website.
Google Analytics is a powerful tool that tracks and analyzes website traffic for informed marketing decisions.
Service URL: policies.google.com (opens in a new window)
You can find more information in our Cookie Policy and Privacy Policy.