yasaklibot
The yasaklibot is a web crawler operated by the Yasaklı project, an open‑source initiative (GitHub: github.com/yasakli) focused on monitoring internet censorship in Turkey. Its primary purpose is to systematically scan websites to detect whether they are blocked or filtered by Turkish ISPs, aggregating this data into a public transparency database. The bot is maintained by a community of researchers and volunteers, and its findings are used to generate reports on censorship patterns.
The crawler uses Python’s Scrapy framework and issues HTTP/HTTPS requests from a rotating pool of IP addresses primarily hosted on Hetzner and DigitalOcean (documented in the project’s GitHub README). It sends approximately 10‑15 requests per minute per IP to avoid overwhelming target servers. The bot checks for HTTP status codes, redirects (especially to error pages served by ISPs), and DNS resolution failures—comparing results between a Turkish‑based IP and a foreign exit node. It does not follow links aggressively; instead, it focuses on a predefined list of news and social media URLs curated from the project’s own database.
According to the yasakli/crawler repository on GitHub, the bot fully respects robots.txt Disallow directives by default. The source code includes a middleware that loads and parses robots.txt before any request, and if a path is disallowed, the request is dropped without a retry. This behavior has been confirmed by site administrators who have inspected the crawler’s logs.
The default User‑Agent string is yasaklibot/1.0 (+https://github.com/yasakli/crawler). The bot also sets the From header to [email protected] and does not spoof common browser strings. Its requests often have a low timeout (5 seconds) and lack Accept‑Language headers, which distinguishes it from typical browsers. Additionally, the bot includes a custom X‑Crawler‑Name header set to yasaklibot.
Collected data—including blocked URLs, redirect chains, and ISP error pages—is published on the yasakli.com website as an interactive map and downloadable CSV files. This information is used for academic research, journalism, and public awareness about internet censorship in Turkey. No personal data is stored, and the project explicitly states it does not sell or monetize the data.
Although yasaklibot is non‑malicious, it is rate‑limited because its scanning of hundreds of URLs per hour can still degrade performance on small websites, especially when combined with other automated agents. Threshold‑based blocking (e.g., 20 requests per IP per minute) is recommended to protect server resources while allowing the bot to complete its censorship‑detection mission.
Similar Threats
Free Bot Analysis
Find out exactly how much of your traffic is automated — and which bots are draining your bandwidth and skewing your analytics.
Run Free Bot Scan →No credit card required · Results in minutes
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.
Stay up to date with the latest from Boteraser.
We use cookies to improve your experience on our site. By using our site, you consent to cookies.
Manage your cookie preferences below:
Essential cookies enable basic functions and are necessary for the proper function of the website.
CloudFlare provides web performance and security solutions, enhancing site speed and protecting against threats.
Service URL: developers.cloudflare.com (opens in a new window)
These cookies are needed for adding comments on this website.
These cookies are used for managing login functionality on this website.
Statistics cookies collect information anonymously. This information helps us understand how visitors use our website.
Google Analytics is a powerful tool that tracks and analyzes website traffic for informed marketing decisions.
Service URL: policies.google.com (opens in a new window)
You can find more information in our Cookie Policy and Privacy Policy.