forschungsportal
Forschungsportal is a web crawler operated by the German National Library of Science and Technology (TIB) – Leibniz Information Centre for Science and Technology, as part of its Forschungsportal service (research-portal.de). The bot systematically harvests metadata and full-text documents from academic repositories, university publication servers, and open-access journals to build a central index of scientific literature for German researchers and international discovery platforms. Its primary output feeds into the TIB’s own search engine, which currently aggregates over 80 million bibliographic records from more than 5,000 data sources worldwide.
According to TIB’s official crawler policy (published at tib.eu/en/crawler-policy), the forschungsportal bot performs scheduled, batch-oriented crawling with a default request rate of one request per 10 seconds per host to avoid overloading servers. The crawler follows HTTP Last-Modified and ETag headers to minimize redundant downloads, and supports HTTPS exclusively. IP ranges are primarily drawn from TIB’s own subnet 194.95.0.0/16 and occasionally from academic networks in the DFN (German Research Network) allocation. The bot uses a fixed crawl depth of 3 per seed URL and respects robots.txt directives before any request. It identifies itself via the User-Agent string forschungsportal/1.0 and also sends a From header containing the email [email protected] for direct contact.
TIB officially confirms that forschungsportal fully honors Disallow, Crawl-Delay, and Allow directives as specified in the Robots Exclusion Standard. The bot checks robots.txt before every crawl session and caches the content for up to 24 hours, re‑fetching if a change is detected. This behavior is documented in TIB’s open‑source crawler framework available on GitHub (tib‑hannover/forschungsportal‑crawler), where the compliance logic is implemented in Python using the urllib.robotparser module.
The definitive User‑Agent string is forschungsportal/1.0, sometimes appended with a version suffix like forschungsportal/1.0 (TIB; +https://www.tib.eu/en/crawler). Additional fingerprints include the From header [email protected], a fixed Accept‑Encoding: gzip, deflate header, and a Accept‑Language: de‑DE,en;q=0.9. The bot does not set any Referer header and uses a persistent connection pool with up to 4 concurrent connections per host. Server logs can identify it by the reverse DNS lookup resolving to *.tib.eu or *.han‑tib.de.
The collected metadata and full‑text documents are used exclusively for academic discovery and indexing within the Forschungsportal service, which provides a single search interface across German institutional repositories. TIB also shares aggregated, anonymized usage statistics with partner institutions under the Open Access 2020 initiative. No data is sold or used for commercial AI training; the corpus feeds into the Bielefeld Academic Search Engine (BASE) and the OpenAIRE European research graph.
Web administrators should enforce a rate limit of 10 requests per minute from the TIB subnet because the bot’s default 10‑second interval can still saturate low‑capacity servers during initial deep crawls. This threshold-based blocking is recommended solely to protect server stability, not to block the bot outright – the policy rationale is asset protection, not adversarial defense.
Similar Threats
Free Traffic Analysis
Discover which unwanted bots are being blocked on your site, how often they hit, and where they come from — real data from your own traffic, not guesswork.
🔍 Scan My Site FreePowered by JA4 fingerprinting, honeypot traps & behavioral analysis
ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.
Stay up to date with the latest from Boteraser.
We use cookies to improve your experience on our site. By using our site, you consent to cookies.
Manage your cookie preferences below:
Essential cookies enable basic functions and are necessary for the proper function of the website.
CloudFlare provides web performance and security solutions, enhancing site speed and protecting against threats.
Service URL: developers.cloudflare.com (opens in a new window)
These cookies are needed for adding comments on this website.
These cookies are used for managing login functionality on this website.
Statistics cookies collect information anonymously. This information helps us understand how visitors use our website.
Google Analytics is a powerful tool that tracks and analyzes website traffic for informed marketing decisions.
Service URL: policies.google.com (opens in a new window)
You can find more information in our Cookie Policy and Privacy Policy.