io.github.cyanheads/biorxiv-mcp-server

biorxiv-mcp-server

Search and retrieve bioRxiv and medRxiv preprints — by DOI, date interval, or keyword — via MCP.

0.3.1
Version
remote + npm
Transport
6
Tools

Security review

Review passed

Reviewed 38m ago.

  • tools: 6 tools scanned
  • metadata: scanned
  • packages: 2 checked

No findings.

Tools (6)

  • biorxiv_list_categories

    List valid subject category strings for bioRxiv and medRxiv — the categories the listing API actually filters on. Use these strings as the `category` filter in biorxiv_list_recent to narrow results to a specific field; case, and "_" or "-" in place of a space, do not matter there. Run this tool before filtering to get the current valid values.

  • biorxiv_list_recent

    List preprints posted or revised within a date interval, optionally scoped to one server or a subject category. Returns 30 preprints per page (fixed by the API); pass `cursor` as an integer offset (0, 30, 60, …) to step through additional pages. Abstracts are omitted by default to keep the page small — pass include_abstract: true for the whole page, or call biorxiv_get_preprint (up to 10 DOIs per call) for a few. When server="both" (default), per-server pagination state is returned separately — use each server's `cursor` field for independent advancement. One server failing under server="both" does not abort the call: the other server's page is still returned and the failed one is named in `failed[]`, marking the result set as partial rather than complete. Every attempted server failing is a different case and does abort the call, with a retryable upstream_unavailable (or rate_limited) error — an empty page would otherwise be indistinguishable from an interval that genuinely holds noth

  • biorxiv_get_preprint

    Fetch full metadata, abstract, all revision history, JATS XML full-text links, and published-journal DOI for one or more preprints by DOI. Each DOI returns all revisions in one response. When server="both" (default), each DOI is checked against both bioRxiv and medRxiv; the response includes which server the preprint was found on. Failed lookups are reported per-DOI in failed[] rather than aborting the batch, each carrying a reason (not_found, invalid_doi_format, upstream_unavailable, rate_limited) and a retryable flag; a rate_limited entry also carries the wait in seconds the origin asked for. DOIs must match the pattern 10.NNNN/…; a doi.org or article URL, a doi: label, and a trailing vN / .full suffix are stripped first, and results report the bare DOI.

  • biorxiv_get_published_version

    Resolve a preprint DOI to its full journal publication record — journal DOI, journal name, published date, and corresponding author details. Use when the preprint's `publishedJournalDoi` field from biorxiv_get_preprint is present and you need the full crosswalk metadata. bioRxiv and medRxiv share their DOI prefixes, so server="both" (the default) checks both in parallel and the response reports which server answered. Works for 10.1101/ and 10.64898/ DOIs alike; when the crosswalk holds no record for a published preprint, the journal DOI still comes back from the preprint's own record, without journal name or date, and a notice says so. Returns a not-found error only when no server holds the preprint or it lists no journal version at all.

  • biorxiv_search_preprints

    Search preprints by keyword and/or author using EuropePMC for relevance ranking, then enrich matching DOIs with full bioRxiv/medRxiv metadata. Provide a keyword query, an author name, or both — author maps to an EuropePMC AUTH: field query and is ANDed with the keyword query. Covers both servers by default. EuropePMC indexes new preprints within 1–2 days of posting; for preprints posted within the last day, prefer biorxiv_list_recent. Abstracts are included by default; include_abstract: false omits them from every result for a response about a third the size, and biorxiv_get_preprint returns the abstract for up to 10 DOIs per call. A EuropePMC rate limit (HTTP 429) fails the call with a retryable rate_limited error carrying the wait in seconds — a rate-limited metadata enrichment does not, and instead marks the affected record enrichment_error: "rate_limited".

  • biorxiv_get_fulltext

    Retrieve a preprint's full text as best-effort Markdown, extracted from its rendered HTML article page. Reads the latest version unless one is requested (the version input, or a vN suffix on the DOI), confirms it via the details API, then fetches and extracts the body — abstract, sections, and references. bioRxiv and medRxiv share the 10.1101/ DOI prefix, so server="both" (the default) resolves the DOI against both in parallel and the response reports which server answered. This is HTML-to-Markdown extraction, not structured JATS: section structure is approximate and not guaranteed. Long articles exceed a single response, so use offset and limit to page through them (the response reports totalChars, remainingChars, and hasMore); paging is cheap because the extracted article is cached per version for an hour after the first read, so only the first chunk pays for a fetch. Not every preprint has an extractable HTML page — some are PDF-only and some origins block programmatic access — in w