Agent Skills

openalex-mcp-server

Access the OpenAlex academic research catalog - 270M+ publications through MCP. STDIO & Streamable HTTP.

Install

npx -y @cyanheads/openalex-mcp-server
  • OPENALEX_API_KEYoptional — OpenAlex account API key, sent as api_key= - optional (free from https://openalex.org/settings/api). Without it, anonymous rate limits apply.
  • OPENALEX_MAILTOoptional — Optional email sent as mailto= to identify yourself to OpenAlex (the polite pool). Separate from the API key.
  • OPENALEX_BASE_URLoptional — OpenAlex API base URL.
  • MCP_LOG_LEVELoptional — Sets the minimum log level for output (e.g., 'debug', 'info', 'warn').
README.md

@cyanheads/openalex-mcp-server

Access the OpenAlex academic research catalog - 270M+ publications through MCP. STDIO & Streamable HTTP.

5 Tools • 2 Prompts

Version License Docker MCP SDK npm TypeScript Bun

Install in Claude Desktop Install in Cursor Install in VS Code

Framework


Overview

Scholarly catalog data from OpenAlex — 270M+ works, 90M+ authors, 100K+ sources, plus institutions, topics, keywords, publishers, and funders. Search, filter, and aggregate across all eight entity types, resolve ambiguous names to canonical IDs, and walk the citation graph one hop at a time. Runs as a stdio process, a local Streamable HTTP server, or the public hosted endpoint above.

Tools

Tool Description
openalex_search_entities Search, filter, sort, or retrieve by ID across all 8 entity types
openalex_analyze_trends Group-by aggregation for trend and distribution analysis
openalex_resolve_name Resolve a name or an identifier (DOI, ORCID, ROR, PMID, ISSN, OpenAlex ID) to an OpenAlex ID
openalex_get_citation_graph Walk the citation graph one hop from a seed work: cites, cited_by, or related_to
openalex_describe_fields List valid filter, group_by, and select field names for an entity type

Prompts

Prompt Description
openalex_literature_review Guides a systematic literature search: formulate query, search, filter, analyze citation network, synthesize findings
openalex_research_landscape Analyzes the research landscape for a topic: volume trends, top authors/institutions, open access rates, funding sources

Capability reference

openalex_search_entities tool

  • Retrieve a single entity by ID — OpenAlex ID, DOI, ORCID, ROR, PMID, ISSN, or PMCID (bare or URL form), or a keyword by its slug or keyword URL. id takes precedence: search parameters passed alongside it are dropped, and the response names which ones. A PMCID resolves nothing (OpenAlex indexes none) — use the work's PMID or DOI instead
  • Keyword search (boolean operators, quoted phrases, wildcards, fuzzy match) plus exact and semantic search modes — semantic ranks a query-dependent candidate set at ~1 req/sec (up to 50 per page), and its meta.count reports the candidate count rather than a match total
  • Rich filter syntax: AND across fields, OR within a field (|), NOT (!), ranges, comparisons; a comma inside a filter value is rejected (use |, or a .search filter for free text)
  • select returns a curated per-entity-type default unless overridden, or ["*"] for the full record; invalid field names error with the valid set
  • Cursor pagination for keyword and exact search, up to 100 results per page (default 25); semantic search walks its candidates with page (1-based) instead, and mixing either knob with the wrong search_mode is rejected before the upstream call
  • sample (up to 100, single page only — neither cursor nor page applies) plus a deterministic seed for reproducible random sampling; keyword and exact modes only, and never alongside sort — OpenAlex does not sample a semantic search and refuses a sorted sample, so both combinations are rejected before the upstream call
  • Each response holds to 64,000 bytes per surface (structuredContent, and content[] with its trailer) unless the least it can return is larger, which it flags as over_budget: a page that would run over keeps its leading records whole and names the rest in omitted with the one call that returns them (as a set, in OpenAlex's order), without disturbing next_cursor; a record too large on its own comes back with its largest arrays windowed, filled in turn so each shows what fits and labeled with how much of the array it holds, and slice: {field, offset} on an id lookup pages any array to the end
  • A list record carrying exactly 100 authorships — the most OpenAlex returns in a list response — is flagged as possibly capped, with the unbilled id + slice call that fetches the rest
  • display_name is nullable for untitled records; every call reports OpenAlex daily-budget cost and remaining balance

openalex_analyze_trends tool

  • Group any supported field for trend, distribution, or comparative analysis; combine with filters to scope the population before aggregation
  • Up to 200 groups per page (default). order: "count" (default) returns the top-N by count with no further pages; order: "key" enumerates all distinct values key-ascending with cursor pagination
  • include_unknown (default false) adds a group for entities with no value for the grouped field, flagged is_unknown: true and labeled as unknown in the text output. OpenAlex keys it -111/-111.0 (numeric fields), unknown (text fields and order: "key"), or an ID ending in /unknown; the key is a sentinel, not a filter value. Boolean fields have no such group — a missing value counts as false
  • Not every field is groupable — raw date fields, .search operators, from_*/to_* range modifiers, decimal scores such as fwci, display_name, and external-ID fields such as doi are rejected; openalex_describe_fields(entity_type, "group_by") lists the fields that group
  • A works group_by on authorships.countries says in its notice that OpenAlex counts only each work's first 100 authorships for that field, and names authorships.institutions.country_code as the field that counts all of them
  • Reports OpenAlex daily-budget cost and remaining balance — aggregation is priced far below paging the same entities

openalex_resolve_name tool

  • A name or partial name runs an autocomplete search: up to 10 matches with disambiguation hints (last institution, host organization, place, etc.)
  • An identifier — OpenAlex ID, DOI, ORCID, ROR, PMID, ISSN, or keyword URL, bare or in URL form — resolves directly to the one record it addresses; no entity_type needed, since the identifier determines its own. A PMCID is recognized but resolves nothing — OpenAlex indexes none
  • filters narrows autocomplete only; on an identifier lookup they're ignored and named in a notice
  • With entity_type set, OpenAlex autocomplete fails on a name over 1,000 characters; that failure is reported once as query_too_long with advice to shorten the name, not retried as an outage
  • Reports OpenAlex daily-budget cost and remaining balance

openalex_get_citation_graph tool

  • direction sets the edge: cites (works citing the seed), cited_by (the seed's own reference list), related_to (OpenAlex's algorithmic related works, ~8-30 typical, may be empty)
  • seed_id accepts an OpenAlex ID, DOI, or PMID (PMCID recognized but resolves nothing); validated against a live lookup first, so a non-existent seed fails as NotFound rather than returning an empty graph
  • Stacks with filters/sort/select to narrow the graph; filters cannot set cites/cited_by/related_to, nor an alias of one such as cited_works — those keys are reserved for direction
  • Cursor pagination, up to 100 results per page (default 25)
  • The same 64,000-byte response budget and 100-authorship cap disclosure as openalex_search_entities; works cut from a page are named in omitted with the openalex_search_entities call that returns them, in OpenAlex's order
  • Reports OpenAlex daily-budget cost, covering both the seed-validation lookup and the graph page, plus remaining balance

openalex_describe_fields tool

  • Lists every valid field name for an entity type + context (filter, group_by, select) — the complete pool, never truncated
  • group_by is the filter set minus what OpenAlex rejects as an aggregation key: raw date fields, .search/.search.exact operators, from_*/to_* range modifiers, and a per-entity-type set found by grouping every listed field against the live API — decimal scores such as fwci, several source year fields, display_name, and external-ID fields among them
  • Optional query reorders results by name similarity without dropping any field — a nested value's parent object stays reachable further down the list
  • Backed by a generated field catalog — no live API calls

openalex_literature_review prompt

  • Arguments: topic required; scope (narrow / broad) optional, defaults to narrow
  • Returns one user message walking a 6-step workflow: resolve entities, search literature, identify key papers, trace citations, analyze the landscape, synthesize findings
  • scope changes the search step: narrow favors exact search with tight topic filters; broad adds semantic search across multiple related topic IDs

openalex_research_landscape prompt

  • Arguments: topic required
  • Returns one user message walking a 7-step quantitative workflow: resolve the topic ID, volume trends, top contributors (institutions/countries/journals), open access rate, funding sources, most-cited works, emerging fronts
  • The funding step groups by awards.funder_id (resolve names via openalex_resolve_name) or awards.funder_display_name for readable labels in a single hop

Features

Built on @cyanheads/mcp-ts-core: stdio and Streamable HTTP transports, pluggable auth (none / jwt / oauth), swappable storage (in-memory, filesystem, Supabase, Cloudflare KV/R2/D1), structured logging with optional OpenTelemetry tracing.

OpenAlex-specific:

  • Typed API client with automatic ID normalization (DOI, ORCID, ROR, PMID, PMCID, ISSN, OpenAlex and PubMed/PubMed Central URLs); a PMCID normalizes but resolves nothing since OpenAlex indexes none
  • Keyless by default — an optional API key raises rate and daily-budget limits, and an optional mailto identifies the caller to OpenAlex's polite pool
  • HTTP status codes mapped to specific MCP error classes (400 → InvalidParams, 422 → ValidationError, 429 → RateLimited) with upstream messages surfaced
  • Timeout-aware request retries and cancellation support via AbortSignal

Agent-friendly output:

  • Provenance — every API-calling tool reports OpenAlex daily-budget cost, remaining balance, and reset time (budget.costUsd, remainingUsd, resetsInSeconds)
  • Effective-query echo — search, trends, and citation-graph responses echo the criteria that actually ran, so an empty result is diagnosable without re-reading the request
  • Discriminated output contracts — typed error reasons (entity_not_found, upstream_budget_exhausted, semantic_per_page_cap, reserved_filter_key, and more) each carrying an explicit recovery hint
  • Response shaping — abstracts are reconstructed from OpenAlex's inverted-index encoding into plaintext, and display_name stays null for untitled or paratext records instead of being backfilled
  • Clean provider text — OpenAlex passes titles, abstracts, and names through unsanitized; HTML entities are decoded against the full WHATWG table, comments and HTML/JATS/MathML formatting tags are removed, CDATA wrappers and IOP's <?CDATA …?> math markers are unwrapped with their TeX kept, and a double-escaped line break (a literal \n) becomes a space; literal text such as A < B and TeX such as \nu is kept. Identifier and URL fields (id, doi, ids, orcid, ror, issn, *_url, lineages) are returned exactly as OpenAlex stores them, so each one resolves when passed back. The text output escapes that text for Markdown, so a stray *, <tag>, or [x](y) renders as written; IDs, and URLs that render as links, stay byte-identical

Getting started

Public Hosted Instance

A public instance is available at https://openalex.caseyjhand.com/mcp — no installation required. Point any MCP client at it via Streamable HTTP:

{
  "mcpServers": {
    "openalex-mcp-server": {
      "type": "streamable-http",
      "url": "https://openalex.caseyjhand.com/mcp"
    }
  }
}

Self-Hosted / Local

Add the following to your MCP client configuration file.

{
  "mcpServers": {
    "openalex-mcp-server": {
      "type": "stdio",
      "command": "bunx",
      "args": ["@cyanheads/openalex-mcp-server@latest"],
      "env": {
        "MCP_TRANSPORT_TYPE": "stdio",
        "MCP_LOG_LEVEL": "info",
        "OPENALEX_API_KEY": "your-api-key"
      }
    }
  }
}

Or with npx (no Bun required):

{
  "mcpServers": {
    "openalex-mcp-server": {
      "type": "stdio",
      "command": "npx",
      "args": ["-y", "@cyanheads/openalex-mcp-server@latest"],
      "env": {
        "MCP_TRANSPORT_TYPE": "stdio",
        "MCP_LOG_LEVEL": "info",
        "OPENALEX_API_KEY": "your-api-key"
      }
    }
  }
}

Or with Docker:

{
  "mcpServers": {
    "openalex-mcp-server": {
      "type": "stdio",
      "command": "docker",
      "args": [
        "run", "-i", "--rm",
        "-e", "MCP_TRANSPORT_TYPE=stdio",
        "-e", "OPENALEX_API_KEY=your-api-key",
        "ghcr.io/cyanheads/openalex-mcp-server:latest"
      ]
    }
  }
}

For Streamable HTTP, set the transport and start the server:

MCP_TRANSPORT_TYPE=http MCP_HTTP_PORT=3010 OPENALEX_API_KEY=... bun run start:http
# Server listens at http://localhost:3010/mcp

OPENALEX_API_KEY is optional — set it to a free OpenAlex account key for keyed rate limits and budget under OpenAlex's usage-based pricing, or omit it for anonymous access. Set OPENALEX_MAILTO to an email if you want to identify yourself to OpenAlex (the polite pool).

Prerequisites

Installation

  1. Clone the repository:
git clone https://github.com/cyanheads/openalex-mcp-server.git
  1. Navigate into the directory:
cd openalex-mcp-server
  1. Install dependencies:
bun install
  1. Configure environment:
cp .env.example .env
# edit .env and set required vars

Configuration

Variable Description Default
MCP_TRANSPORT_TYPE Transport: stdio or http. stdio
MCP_HTTP_PORT Port for HTTP server. 3010
MCP_SESSION_MODE HTTP session mode: stateless, stateful, or auto (resolves to stateful). The server declares stateless in code; an explicit value overrides it. stateless
MCP_AUTH_MODE Auth mode: none, jwt, or oauth. none
MCP_ALLOWED_ORIGINS Comma-separated allow-list of browser Origin headers for HTTP transport. Unset = loopback-only; set to * to disable. loopback only
MCP_LOG_LEVEL Log level (RFC 5424). debug
LOGS_DIR Directory for log files (Node.js only). <project-root>/logs
STORAGE_PROVIDER_TYPE Storage backend. in-memory
OPENALEX_API_KEY OpenAlex account API key, sent upstream as api_key= (free from openalex.org/settings/api). Without it, anonymous rate limits apply. —
OPENALEX_MAILTO Email sent upstream as mailto= to identify yourself to OpenAlex (the "polite pool"); a courtesy identifier, separate from the API key. —
OPENALEX_BASE_URL OpenAlex API base URL. https://api.openalex.org
OTEL_ENABLED Enable OpenTelemetry instrumentation (spans, metrics, completion logs). false

See .env.example for the full list of optional overrides.

Running the server

Local development

  • Build and run:

    # One-time build
    bun run rebuild
    
    # Run the built server
    bun run start:stdio
    # or
    bun run start:http
    
  • Run checks and tests:

    bun run devcheck   # Lints, formats, type-checks
    bun run test       # Runs the test suite
    

Docker

docker build -t openalex-mcp-server .
docker run --rm -e OPENALEX_API_KEY=your-key -p 3010:3010 openalex-mcp-server

The Dockerfile defaults to HTTP transport, stateless session mode, and logs to /var/log/openalex-mcp-server. OpenTelemetry peer dependencies are installed by default — build with --build-arg OTEL_ENABLED=false to omit them.

Project structure

Directory Purpose
src/index.ts createApp() entry point — registers tools and prompts.
src/config/ Server-specific environment variable parsing and validation with Zod.
src/mcp-server/tools/definitions/ Tool definitions (*.tool.ts).
src/mcp-server/prompts/definitions/ Prompt definitions (*.prompt.ts).
src/services/openalex/ OpenAlex API client, field catalog, and domain types.
tests/ Unit and integration tests, mirroring the src/ structure.

Development guide

See CLAUDE.md for development guidelines and architectural rules. The short version:

  • Handlers throw, framework catches — no try/catch in tool logic
  • Use ctx.log for logging, ctx.state for storage
  • Wrap OpenAlex responses: validate the raw payload → normalize to a domain type → return the output schema; never fabricate missing fields
  • Always resolve names to IDs via openalex_resolve_name before filtering by entity

Contributing

Issues are welcome. Run checks before submitting:

bun run devcheck
bun run test

License

Apache-2.0 — see LICENSE for details.

Search skills and MCP servers

Fuzzy search across 23,137 skills and servers