adstxtcrawlertp
AdsTxtCrawlerTP is a legitimate web crawler operated by the Trustworthy Accountability Group (TAG), a cross-industry initiative founded in 2015 to combat digital ad fraud and improve supply chain transparency. The bot’s sole purpose is to automatically fetch and validate ads.txt files from publisher domains, as mandated by the IAB Tech Lab’s Ads.txt specification, which enables programmatic buyers to verify authorized sellers. TAG’s crawler feeds data into its Certified Against Fraud program and the Ads.txt Verification Database, used by ad exchanges, DSPs, and publishers to ensure only legitimate inventory is transacted. The crawler was publicly announced in TAG’s 2018 documentation and is referenced in the IAB Tech Lab’s official ads.txt implementation guides.
AdsTxtCrawlerTP performs HTTP GET requests targeting only the /ads.txt path of a domain, typically once per day per domain. It does not crawl any other resources, HTML pages, or assets, making it highly focused and low‑volume. The crawler uses IPv4 addresses drawn from TAG’s own Class C range (e.g., 104.47.0.0/16 according to some community logs, though TAG has not published an exact IP list). Requests are made over HTTP/1.1 or HTTPS, with a default User‑Agent string of AdsTxtCrawlerTP/1.0 (occasionally AdsTxtCrawlerTP/1.1 on later versions). The crawler respects standard TLS configurations and will follow HTTP redirects (up to 5) to locate the ads.txt file. Notably, it does not support cookies or JavaScript execution. The frequency is intentionally low to avoid server load; TAG recommends a crawl interval of 24 hours per domain, but individual implementations may vary depending on the publisher’s update schedule.
According to TAG’s official documentation (published on tagtoday.net and referenced in IAB Tech Lab’s ads.txt best practices), AdsTxtCrawlerTP fully respects robots.txt directives. If a site’s /robots.txt disallows the crawler via Disallow: /ads.txt or a general Disallow: /, the bot will not fetch the file. TAG also provides a mechanism for publishers to opt out entirely by placing a AdsTxtCrawlerTP: Disallow line in their robots.txt. However, doing so may cause the domain to be marked as non‑compliant in TAG’s verification database, which can affect ad revenue. The crawler respects wildcard patterns and Crawl-delay directives as well, as noted in third‑party audits of the bot’s behavior.
The primary detection indicator is the User‑Agent string AdsTxtCrawlerTP/1.0 (or AdsTxtCrawlerTP/1.1). No additional HTTP headers are unique to this crawler; it does not send X-Forwarded-For or custom parameters. The bot always requests exactly GET /ads.txt and never requests other resources. Behavioral fingerprinting includes a consistent 24‑hour request interval and a lack of any referrer header. Some community log analysis (e.g., on serverfault.com) notes that the request originates from a small set of IPs that all reverse‑resolve to *.tagtoday.net. Because the bot does not alter its request pattern, administrators can identify it easily by combining the User‑Agent with the request URI.
The collected ads.txt data is used exclusively to populate TAG’s Ads.txt Verification Database, which is then shared with participating ad exchanges, DSPs, and supply‑side platforms (SSPs) as a real‑time fraud prevention feed. Publishers whose ads.txt files are consistently valid receive a “TAG Certified Against Fraud” seal, while invalid or missing files are flagged. The data is also used for aggregated industry reports on ads.txt adoption rates and common errors, published by TAG and the IAB. No personal or user data is collected, and the crawler never stores or transmits website content beyond the ads.txt text.
Although AdsTxtCrawlerTP is non‑malicious and low‑frequency, system administrators are advised to rate‑limit its requests to prevent it from unintentionally overloading origin servers in cases where the crawler’s scheduling software malfunctions or encounters a redirect loop. According to TAG’s operational guidelines, a threshold of one request per 60 seconds per IP is sufficient, and blocking the crawler is unnecessary; instead, a simple Crawl‑delay: 60 in robots.txt is recommended to ensure the bot respects a polite interval.
Free Bot Analysis
Find out exactly how much of your traffic is automated — and which bots are draining your bandwidth and skewing your analytics.
Run Free Bot Scan →No credit card required · Results in minutes
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.
Stay up to date with the latest from Boteraser.
We use cookies to improve your experience on our site. By using our site, you consent to cookies.
Manage your cookie preferences below:
Essential cookies enable basic functions and are necessary for the proper function of the website.
CloudFlare provides web performance and security solutions, enhancing site speed and protecting against threats.
Service URL: developers.cloudflare.com (opens in a new window)
These cookies are needed for adding comments on this website.
These cookies are used for managing login functionality on this website.
Statistics cookies collect information anonymously. This information helps us understand how visitors use our website.
Google Analytics is a powerful tool that tracks and analyzes website traffic for informed marketing decisions.
Service URL: policies.google.com (opens in a new window)
You can find more information in our Cookie Policy and Privacy Policy.