poggio-citations
Poggio-Citations is a web crawler operated by the MIT Computer Science and Artificial Intelligence Laboratory (CSAIL), specifically associated with the research group of Tomaso Poggio. It systematically collects academic citation data from publicly accessible scholarly articles, preprints, and institutional repositories to build a comprehensive citation graph for research analysis. The bot feeds data into the Poggio Citation Graph database, which supports scientometric studies and AI research on citation networks.
The bot crawls at a moderate rate of approximately 1–2 requests per second, with controlled bursts during initial indexing of new domains. It uses HTTP/1.1 and HTTPS protocols exclusively, and its IP ranges are allocated from the MIT network (18.0.0.0/8 and 128.30.0.0/16). The crawler focuses on HTML pages containing citation metadata, PDF files, and BibTeX files. It follows hyperlinks within academic domains but avoids non-text content such as images or videos. Default crawl depth is limited to 3 levels from seed URLs, and it respects 404 and 410 status codes by ceasing to follow broken links. Requests include a Referer header pointing to the Poggio project homepage.
According to official documentation from the Poggio project, the bot fully respects robots.txt Disallow directives. It also honors Crawl-Delay directives if specified. The crawler fetches robots.txt upon first visit and caches the file for 24 hours.
The primary User-Agent string is Mozilla/5.0 (compatible; Poggio-Citations/1.0; +https://poggio.mit.edu/crawler). Additionally, the bot sends a custom HTTP header X-Poggio-Crawler: 1. Its behavior includes a consistent request interval of 500–1000 milliseconds, which can be used for fingerprinting. The User-Agent field contains the version number, and the bot does not mask its identity.
Collected citation data is used exclusively for academic research purposes, including training graph neural networks for citation prediction, analyzing research trends, and building open-access citation datasets. The project is non-commercial; all data is publicly released under a CC-BY license. The crawler stores only citation metadata (authors, title, journal, DOI, references) and not full article content.
While Poggio-Citations is a legitimate research crawler, it is rate-limited by many academic publishers due to its systematic, high-volume requests that can tax server resources. A threshold-based blocking policy (e.g., blocking if >100 requests per minute from a single IP) is recommended to protect site availability while still allowing the bot to collect essential citation data.
🛡️
Boteraser automatically detects and blocks unwanted bots — protecting your site from scrapers, DDoS bursts, and credential stuffing attacks without slowing down real visitors.
✅ Start Free ProtectionSetup takes under a minute · Free trial available
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.
Stay up to date with the latest from Boteraser.
We use cookies to improve your experience on our site. By using our site, you consent to cookies.
Manage your cookie preferences below:
Essential cookies enable basic functions and are necessary for the proper function of the website.
CloudFlare provides web performance and security solutions, enhancing site speed and protecting against threats.
Service URL: developers.cloudflare.com (opens in a new window)
These cookies are needed for adding comments on this website.
These cookies are used for managing login functionality on this website.
Statistics cookies collect information anonymously. This information helps us understand how visitors use our website.
Google Analytics is a powerful tool that tracks and analyzes website traffic for informed marketing decisions.
Service URL: policies.google.com (opens in a new window)
You can find more information in our Cookie Policy and Privacy Policy.