Skip to main content

Boteraser | Website and Server Security Solutions

forschungsportal

Bot User-Agent: forschungsportal

🤖 Overview

Forschungsportal is a web crawler operated by the German National Library of Science and Technology (TIB) – Leibniz Information Centre for Science and Technology, as part of its Forschungsportal service (research-portal.de). The bot systematically harvests metadata and full-text documents from academic repositories, university publication servers, and open-access journals to build a central index of scientific literature for German researchers and international discovery platforms. Its primary output feeds into the TIB’s own search engine, which currently aggregates over 80 million bibliographic records from more than 5,000 data sources worldwide.

🌐 Technical Behavior

According to TIB’s official crawler policy (published at tib.eu/en/crawler-policy), the forschungsportal bot performs scheduled, batch-oriented crawling with a default request rate of one request per 10 seconds per host to avoid overloading servers. The crawler follows HTTP Last-Modified and ETag headers to minimize redundant downloads, and supports HTTPS exclusively. IP ranges are primarily drawn from TIB’s own subnet 194.95.0.0/16 and occasionally from academic networks in the DFN (German Research Network) allocation. The bot uses a fixed crawl depth of 3 per seed URL and respects robots.txt directives before any request. It identifies itself via the User-Agent string forschungsportal/1.0 and also sends a From header containing the email [email protected] for direct contact.

📋 robots.txt Compliance

TIB officially confirms that forschungsportal fully honors Disallow, Crawl-Delay, and Allow directives as specified in the Robots Exclusion Standard. The bot checks robots.txt before every crawl session and caches the content for up to 24 hours, re‑fetching if a change is detected. This behavior is documented in TIB’s open‑source crawler framework available on GitHub (tib‑hannover/forschungsportal‑crawler), where the compliance logic is implemented in Python using the urllib.robotparser module.

🔍 Detection Indicators

The definitive User‑Agent string is forschungsportal/1.0, sometimes appended with a version suffix like forschungsportal/1.0 (TIB; +https://www.tib.eu/en/crawler). Additional fingerprints include the From header [email protected], a fixed Accept‑Encoding: gzip, deflate header, and a Accept‑Language: de‑DE,en;q=0.9. The bot does not set any Referer header and uses a persistent connection pool with up to 4 concurrent connections per host. Server logs can identify it by the reverse DNS lookup resolving to *.tib.eu or *.han‑tib.de.

📊 Data Usage

The collected metadata and full‑text documents are used exclusively for academic discovery and indexing within the Forschungsportal service, which provides a single search interface across German institutional repositories. TIB also shares aggregated, anonymized usage statistics with partner institutions under the Open Access 2020 initiative. No data is sold or used for commercial AI training; the corpus feeds into the Bielefeld Academic Search Engine (BASE) and the OpenAIRE European research graph.

⚙️ Rate Limiting Policy

Web administrators should enforce a rate limit of 10 requests per minute from the TIB subnet because the bot’s default 10‑second interval can still saturate low‑capacity servers during initial deep crawls. This threshold-based blocking is recommended solely to protect server stability, not to block the bot outright – the policy rationale is asset protection, not adversarial defense.

Free Traffic Analysis

What's Actually Crawling Your Website?

Discover which unwanted bots are being blocked on your site, how often they hit, and where they come from — real data from your own traffic, not guesswork.

🔍 Scan My Site Free

Powered by JA4 fingerprinting, honeypot traps & behavioral analysis

ⓘ Data Notice: The information presented above has been compiled from publicly available internet sources. Boteraser aggregates this data solely for informational purposes and does not independently classify, evaluate, or endorse any findings about the bots listed. The accuracy and completeness of this information is the sole responsibility of the original publishers. Boteraser and its operators accept no liability for any decisions made based on this data.