MCP, CLI, Skills for searching and downloading academic papers from multiple sources like arXiv, PubMed, bioRxiv, etc.
Install
uvx paper-search-mcpPaper Search MCP
A Model Context Protocol (MCP) server for searching and downloading academic papers from multiple sources. The project follows a free-first strategy: prioritize open and public data sources, support optional API keys when they improve stability or coverage, and keep source-specific connectors extensible for advanced users.
Table of Contents
- Overview
- Project Principles
- MCP Authorization Compatibility
- Features
- Source Strategy
- Sci-Hub Notice
- Installation
- Contributing
- Demo
- Star History
- License
- TODO
Overview
paper-search-mcp is a Python-based tool for searching and downloading academic papers from various platforms. It provides tools for searching papers, downloading PDFs, and extracting text, making it ideal for researchers and AI-driven workflows. It can be used as an MCP server (for Claude Desktop and other MCP clients) or as a Claude Code skill with a CLI interface.
Project Principles
- Free-First: Public and open sources are the default roadmap. Paid or restricted sources are not the core direction of this project.
- Optional API Keys: API keys are supported only when they improve stability, rate limits, or metadata quality. The MCP should still be usable without them whenever possible.
- LLM-Friendly Retrieval: Search results should be standardized, deduplicated, and as complete as possible for downstream LLM workflows.
- Source Transparency: Different sources have different strengths. The MCP should make those tradeoffs explicit instead of pretending every source supports full-text retrieval.
MCP Authorization Compatibility
The bundled MCP server supports stdio (the default), sse, and streamable-http. Network transports bind to 127.0.0.1 by default. Transport support does not make the server an OAuth 2.1 protected resource: it does not implement protected-resource metadata, bearer-token validation, scopes, or OAuth authorization responses.
For a remote protected deployment, keep the backend private and put it behind an MCP/HTTP gateway that implements the MCP authorization and discovery requirements, forwarding only authorized requests. A generic reverse proxy alone does not establish MCP OAuth compliance. Do not expose the unauthenticated backend directly to the internet. Native protected-resource support remains tracked in #25.
Features
- Two-Layer Architecture:
- Layer 1 (Unified Tooling): High-level
search_papersfor multi-source concurrent search & deduplication, anddownload_with_fallbackrelying on publisher open access links with sequential fallbacks. - Layer 2 (Platform Connectors): Modular connectors for specific academic platforms (arXiv, PubMed, bioRxiv, Semantic Scholar, etc.) equipped with intelligent DOI extraction via regex text analysis or API fields.
- Layer 1 (Unified Tooling): High-level
- Multi-Source Support: Search and download papers from arXiv, PubMed, bioRxiv, medRxiv, Google Scholar, IACR ePrint Archive, Semantic Scholar, Crossref, OpenAlex, PubMed Central (PMC), CORE, Europe PMC, dblp, OpenAIRE, CiteSeerX, DOAJ, BASE, Zenodo, HAL, SSRN, Unpaywall (DOI lookup), and optional Sci-Hub workflows.
- Opt-in Fast Search: CLI search keeps broad coverage by default. Use
-s fast(OpenAlex, Crossref, arXiv, PubMed, Europe PMC) or-s fastest(OpenAlex and Crossref) when lower latency matters more than coverage. - Standardized Output: Papers are returned in a consistent dictionary format via the
Paperclass. - Free-First Design: Open and public sources are prioritized before any optional commercial or restricted integrations.
- Optional API-Key Enhancement: Sources like Semantic Scholar can work better with a user-provided API key, but are not intended to force paid usage.
- Discovery + Retrieval Workflow: Google Scholar and Crossref can be used for discovery and DOI backfilling, while open repositories and publisher links are used for lawful full-text resolution where available.
- OA-First Fallback Chain:
download_with_fallbacknow follows source-native download → OpenAIRE/CORE/Europe PMC/PMC discovery → Unpaywall DOI resolution → optional Sci-Hub. - MCP Integration: Compatible with MCP clients for LLM context enhancement.
- Extensible Design: Easily add new academic platforms by extending the
academic_platformsmodule.
Source Strategy
The long-term goal is not to depend on a single search engine, but to combine multiple free and public sources with clear roles:
- Open metadata backbone: Crossref, OpenAlex, Semantic Scholar, dblp, CiteSeerX, SSRN, Unpaywall (DOI-centric OA metadata).
- Discipline-specific sources: arXiv, PubMed, PubMed Central, Europe PMC, IACR.
- Open-access full-text sources: arXiv, PMC, CORE, OpenAIRE, DOAJ, BASE, Zenodo, HAL, publisher open-access links.
- Discovery and DOI recovery: Google Scholar can be useful for finding titles, versions, and DOI clues when other public metadata sources are incomplete.
Recommended free-first roadmap:
- Keep current public sources stable.
- Add OpenAlex as a broad free metadata source.
- Add PubMed Central and Europe PMC for stronger biomedical full-text access.
- Add CORE and OpenAIRE for repository-based open-access retrieval.
- Use Google Scholar mainly as a discovery fallback, not as the primary canonical source.
Platform Capability Matrix
This matrix reflects verified live-integration results from functional and end-to-end regression tests in this repository. Columns show the highest capability level observed under normal conditions.
| Platform | Search | Download | Read | Notes |
|---|---|---|---|---|
| arXiv | ✅ | ✅ | ✅ | Open API; reliable |
| PubMed | ✅ | ❌ | ⚠️ info-only | Open API; reliable |
| bioRxiv | ✅ | ✅ | ✅ | Open API; reliable |
| medRxiv | ✅ | ✅ | ✅ | Open API; reliable |
| Google Scholar | ⚠️ | ❌ | ❌ | Upstream bot-detection/rate limits can prevent search; reported as source errors |
| IACR | ✅ | ✅ | ✅ | Open API; reliable |
| Semantic Scholar | ✅ | ✅ (OA) | ✅ (OA) | Works without key (rate-limited); key improves limits; key rejection (403) retried automatically without key |
| Crossref | ✅ | ❌ | ⚠️ info-only | Open API; reliable |
| OpenAlex | ✅ | ❌ | ⚠️ info-only | Open API; free API key improves daily limits |
| PMC | ✅ | ✅ (OA only) | ✅ (OA only) | OA PDFs only; direct download may be blocked by some proxy environments |
| CORE | ✅ | ✅ (record-dependent) | ✅ (record-dependent) | Free key recommended; connector retries with backoff and falls back to key-less on 401/403 |
| Europe PMC | ✅ | ✅ (OA) | ✅ (OA) | OA PDFs only; direct download may be blocked by some proxy environments |
| dblp | ✅ | ❌ | ⚠️ info-only | Open API; reliable |
| OpenAIRE | ✅ | ❌ | ❌ | Open API; retries 3× with escalating request profiles on transient 403 |
| CiteSeerX | ⚠️ | ✅ (record-dependent) | ⚠️ | API endpoint intermittently unavailable / redirects to web archive |
| DOAJ | ✅ | ⚠️ (URL-dependent) | ⚠️ (URL-dependent) | PDF availability varies by article; free key raises rate limits |
| BASE | ⚠️ | ✅ (record-dependent) | ✅ (record-dependent) | OAI-PMH endpoint requires institutional IP registration; returns empty gracefully otherwise |
| Zenodo | ✅ | ✅ (record-dependent) | ✅ (record-dependent) | Open API; reliable |
| HAL | ✅ | ✅ (record-dependent) | ✅ (record-dependent) | Open API; reliable |
| SSRN | ⚠️ | ⚠️ best-effort | ⚠️ best-effort | 403 bot-detection active; public PDF only |
| Unpaywall | ✅ (DOI lookup) | ❌ | ❌ | Requires PAPER_SEARCH_MCP_UNPAYWALL_EMAIL |
| Sci-Hub (optional) | ⚠️ fallback-only | ✅ | ❌ | Optional; unstable mirrors; user responsibility |
| IEEE Xplore 🔑 | 🚧 skeleton | 🚧 skeleton | 🚧 skeleton | Requires PAPER_SEARCH_MCP_IEEE_API_KEY to activate |
| ACM DL | ✅ (Crossref metadata) | ⚠️ | ⚠️ | Keyless search; direct PDF/read may be blocked by browser challenges; use OA fallback |
✅ = reliable in live tests. ⚠️ = works but subject to upstream instability or access restrictions. ❌ = not supported. 🔑 = key required. 🚧 = skeleton only.
Credential & API Key Requirements
All keys are optional unless noted. Configure them in ~/.config/paper-search-mcp/.env (preferred) or as shell exports.
| Environment Variable | Provider | Required? | How to obtain |
|---|---|---|---|
PAPER_SEARCH_MCP_UNPAYWALL_EMAIL |
Unpaywall | Yes (Unpaywall disabled without it) | Any valid email; register at unpaywall.org |
PAPER_SEARCH_MCP_CORE_API_KEY |
CORE | Recommended | Free at core.ac.uk/services/api |
PAPER_SEARCH_MCP_SEMANTIC_SCHOLAR_API_KEY |
Semantic Scholar | Optional | Free at semanticscholar.org — improves rate limits |
PAPER_SEARCH_MCP_OPENALEX_API_KEY |
OpenAlex | Optional | Free at openalex.org/settings/api; increases the keyless daily budget 10x |
PAPER_SEARCH_MCP_OPENALEX_EMAIL |
OpenAlex | Optional | Contact email used in the OpenAlex User-Agent |
PAPER_SEARCH_MCP_GOOGLE_SCHOLAR_PROXY_URL |
Google Scholar | Optional | Your HTTP/HTTPS proxy URL; does not guarantee access or remove provider limits |
PAPER_SEARCH_MCP_DOAJ_API_KEY |
DOAJ | Optional | Free at doaj.org — raises hourly rate limit |
PAPER_SEARCH_MCP_ZENODO_ACCESS_TOKEN |
Zenodo | Optional | Free at zenodo.org — required for private records |
PAPER_SEARCH_MCP_IEEE_API_KEY |
IEEE Xplore | Required to activate | Free at developer.ieee.org |
All variables follow the PAPER_SEARCH_MCP_<NAME> prefix scheme. Legacy names without the prefix (e.g. CORE_API_KEY, UNPAYWALL_EMAIL) are still supported for backward compatibility.
Known Upstream Limitations
Some search failures are caused by external provider instability, not by bugs in this project:
| Source | Symptom | Cause | Workaround |
|---|---|---|---|
| Google Scholar | Source error for HTTP/network failures, CAPTCHA, or persistent consent pages | Upstream rate limits, access checks, or connectivity | Reduce request frequency or use another public source; CAPTCHA is not solved automatically |
| Semantic Scholar | 429 rate-limited responses | Anonymous access rate limit | Set PAPER_SEARCH_MCP_SEMANTIC_SCHOLAR_API_KEY; if key is rejected (403) connector automatically retries without key |
| OpenAlex | 403/429 or daily quota errors | Anonymous access daily limit | Set PAPER_SEARCH_MCP_OPENALEX_API_KEY |
| CORE | 500 / timeout errors | Unauthenticated rate limiting | Set PAPER_SEARCH_MCP_CORE_API_KEY (free); connector retries with exponential backoff and falls back to key-less on 401/403 |
| OpenAIRE | Transient 403 responses | IP-based session rate limiting | Connector retries 3× per profile, escalating: plain session → XML Accept header → raw requests.get with Mozilla UA |
| CiteSeerX | 404 via web archive redirect | PSU endpoint intermittently redirects to archive | No workaround; connector returns empty gracefully |
| BASE | Search returns 0 results | OAI-PMH endpoint requires institutional IP registration | Register at base-search.net for API access; connector returns empty gracefully otherwise |
| SSRN | HTTP 403 | Bot-detection (Cloudflare) | No workaround; connector tries two endpoints and returns a clear message on failure |
| PMC / Europe PMC | PDF download ProxyError | Local proxy blocking direct HTTPS PDF download | Disable proxy or use download_with_fallback instead |
| Unpaywall | Skipped entirely | UNPAYWALL_EMAIL env var not set |
Set PAPER_SEARCH_MCP_UNPAYWALL_EMAIL in ~/.config/paper-search-mcp/.env |
Google Scholar failures are exposed in errors.google_scholar by unified MCP
and CLI search, while successful sources still return their papers. Direct
Scholar searches raise GoogleScholarSearchError for those failures. A normal
empty result page still returns an empty list. If a later page fails, the source
is reported as failed rather than returning its earlier pages as complete.
Existing tool/deadline timeout behavior is unchanged. These diagnostics do not
guarantee access to Scholar or a fixed number of queries per session.
Optional Paid Platform Connectors (Phase 3)
IEEE Xplore is an opt-in skeleton, disabled until its API key is configured.
ACM Digital Library search is keyless and enabled by default, using Crossref metadata restricted to ACM DOI prefix 10.1145.
| Platform | Env Var | Status |
|---|---|---|
| IEEE Xplore | PAPER_SEARCH_MCP_IEEE_API_KEY |
🚧 skeleton — search registered, download/read raise NotImplementedError |
| ACM Digital Library | None | Crossref-backed search; PDF download/read depend on publisher access |
How to enable:
export PAPER_SEARCH_MCP_IEEE_API_KEY=<your_ieee_key> # free key at https://developer.ieee.org/
With an IEEE key, ieee and its tools are registered at startup. ACM (acm, search_acm, download_acm, and read_acm_paper) is always available. Legacy PAPER_SEARCH_MCP_ACM_API_KEY / ACM_API_KEY settings are no longer used and can be removed.
ACM downloads require an ACM DOI such as 10.1145/.... Publisher browser challenges may block scripted access; use download_with_fallback(source="acm", paper_id="10.1145/...", doi="10.1145/...") to try open repositories. read_acm_paper downloads a PDF into save_path before extracting text and can overwrite that file.
Free Source Expansion (Phase 4)
Three additional free-source connectors are now integrated into the MCP server:
zenodo: Official Zenodo REST API connector (search + record-dependent PDF/read support).hal: HAL public API connector (search + record-dependent PDF/read support).ssrn: Discovery-first connector with hardened parser and best-effort download/read when a direct public PDF link is available.unpaywall: DOI-centric OA metadata source for standalone lookup (search_unpaywall) and fallback URL resolution.
SSRN integration remains compliance-first: it only attempts direct public PDF links exposed by SSRN pages. If login/restricted delivery is required, the connector returns a clear message instead of bypassing access controls.
Sci-Hub Notice
Sci-Hub support can remain available as an optional connector for users who explicitly choose to enable it, but it should not be treated as the default or recommended full-text path.
download_with_fallbackleaves Sci-Hub disabled by default. Passuse_scihub=trueonly when you explicitly choose to use it.- Availability is unstable and mirrors change frequently.
- Legal and policy risks vary by jurisdiction.
- README and tool descriptions should clearly state that users are responsible for enabling and using it.
- Open-access and publisher-permitted sources should be tried first whenever possible.
Installation
Choose the method that best fits your workflow. All methods support the same optional API keys.
Claude Code (Skill) — recommended for Claude Code users
Install as a Claude Code skill instead of an MCP server. This gives Claude automatic access to paper search when you mention finding papers, academic literature, etc. — no MCP configuration needed.
Prerequisites: uv and Claude Code.
Step 1 — Install the CLI:
uv tool install paper-search-mcp
Step 2 — Install the skill:
mkdir -p ~/.claude/skills/paper-search
curl -fsSL https://raw.githubusercontent.com/openags/paper-search-mcp/main/claude-code/SKILL.md \
-o ~/.claude/skills/paper-search/SKILL.md
Step 3 (optional) — Configure API keys:
Create ~/.config/paper-search-mcp/.env for optional API keys (see Environment Variables).
That's it. Next time you start Claude Code, just ask it to find papers — the skill activates automatically. For example:
- "Find me recent papers on CRISPR base editing"
- "Search arxiv and semantic scholar for transformer attention mechanisms"
- "Download the PDF for arxiv paper 2106.12345"
The skill uses a CLI (paper-search) that wraps the same library as the MCP server, outputting JSON for search/download and plain text for read.
Choose sources explicitly when latency matters:
paper-search search "gender imbalance neuroscience references" -s fast -n 3
paper-search search "gender imbalance neuroscience references" -s fastest -n 3
paper-search search "gender imbalance neuroscience references" -s all -n 3
paper-search download semantic DOI:10.1038/s41593-020-0658-y -o ./downloads
The CLI keeps -s all as its default broad source set. --exhaustive is an
accepted compatibility no-op because broad search is already the default.
Explicit -s selections always take precedence, including -s fast and
-s fastest. The broad set is unchanged;
optional paid sources are never added to these presets.
-s fast selects OpenAlex, Crossref, arXiv, PubMed, and Europe PMC. A nonblank
PAPER_SEARCH_MCP_SEMANTIC_SCHOLAR_API_KEY (or legacy
SEMANTIC_SCHOLAR_API_KEY) also includes Semantic Scholar. -s fastest selects
only OpenAlex and Crossref. Presets are latency-oriented choices, not guarantees
of response time or exhaustive literature coverage. Anonymous Semantic Scholar
429 responses fail immediately; authenticated requests retain bounded retries.
Only selected searchers are constructed. Source lists and presets are honored
exactly: a DOI in the query never adds another source. Use -s unpaywall for a
DOI lookup, or include unpaywall explicitly in a comma-separated source list.
The broad all preset already includes it. MCP server search defaults are
unchanged.
Sort search results by citation count or publication date:
paper-search search "transformer attention" --sources arxiv,semantic --sort citations
paper-search search "CRISPR" --sort date
--sort relevance (the default) preserves the existing order: selected sources
in order, with each source's returned order unchanged. It does not compute a
cross-source relevance score. --sort citations orders highest counts first;
--sort date orders newest publication dates first. Sorting is client-side,
after deduplication, and only covers retrieved results (--max-results is per
source); it does not change source queries or search the full source collection
for its most-cited/newest papers. Ties retain their original order. Missing or
invalid values sort last; numeric citation strings are supported. Dates accept
ISO dates/timestamps, with naive timestamps and date-only values treated as UTC.
The default JSON output and selected sources are unchanged.
Skill ZIP uploads and other Claude runtimes
The steps above install a local Claude Code skill at ~/.claude/skills/paper-search/SKILL.md, following the Claude Code skill layout. They do not require a ZIP upload.
GitHub's Download ZIP contains the entire repository, with the skill nested under paper-search-mcp-main/claude-code/. It is not a standalone skill archive. If an uploader reports that SKILL.md is nested too deeply, check the archive layout first. Anthropic's custom-skill packaging guide specifies a single skill folder at the archive root, with the skill file directly inside it:
paper-search-skill.zip
└── paper-search/
└── SKILL.md
To create that layout for inspection or adaptation, run this from a repository checkout (requires Python 3; overwrites paper-search-skill.zip):
from zipfile import ZIP_DEFLATED, ZipFile
with ZipFile("paper-search-skill.zip", "w", compression=ZIP_DEFLATED) as archive:
archive.write("claude-code/SKILL.md", arcname="paper-search/SKILL.md")
This only packages the instructions; it does not bundle Python dependencies, install paper-search, or configure an MCP connection. The bundled skill expects a runtime that can execute the CLI and reach the academic services. Upload acceptance and execution in Claude web or another runtime have not been validated by this project. Check that runtime's current metadata, package-installation, network-access, and code-execution requirements before adapting the skill. If you want to use an MCP client instead, follow the MCP installation methods below; uploading a skill ZIP does not start or connect an external MCP server.
MCP Server Config file locations (for methods below)
- macOS:
~/Library/Application Support/Claude/claude_desktop_config.json- Windows:
%APPDATA%\Claude\claude_desktop_config.json- Linux:
~/.config/Claude/claude_desktop_config.json
Method 1 — Smithery (one-command, recommended for Claude Desktop)
npx -y @smithery/cli install @openags/paper-search-mcp --client claude
Smithery automatically writes the correct config block for you. No manual JSON editing needed.
Method 2 — uvx (no install, always latest)
uvx runs the package directly from PyPI without a permanent install. Requires uv.
# Install uv (skip if already installed)
curl -LsSf https://astral.sh/uv/install.sh | sh
⚠️ macOS note:
uvxgenerated wrapper scripts rely onrealpath, which is not included in macOS by default. If you see arealpath: command not founderror, either install GNU coreutils (brew install coreutils) or use Method 3 (uv run) instead — it does not have this limitation.
Claude Desktop config:
{
"mcpServers": {
"paper-search-mcp": {
"command": "uvx",
"args": ["paper-search-mcp"],
"env": {
"PAPER_SEARCH_MCP_UNPAYWALL_EMAIL": "your@email.com",
"PAPER_SEARCH_MCP_CORE_API_KEY": "",
"PAPER_SEARCH_MCP_SEMANTIC_SCHOLAR_API_KEY": "",
"PAPER_SEARCH_MCP_ZENODO_ACCESS_TOKEN": "",
"PAPER_SEARCH_MCP_GOOGLE_SCHOLAR_PROXY_URL": "",
"PAPER_SEARCH_MCP_IEEE_API_KEY": ""
}
}
}
}
Method 3 — uv (persistent install)
uv tool install paper-search-mcp
Claude Desktop config:
{
"mcpServers": {
"paper-search-mcp": {
"command": "uv",
"args": ["tool", "run", "paper-search-mcp"],
"env": {
"PAPER_SEARCH_MCP_UNPAYWALL_EMAIL": "your@email.com",
"PAPER_SEARCH_MCP_CORE_API_KEY": "",
"PAPER_SEARCH_MCP_SEMANTIC_SCHOLAR_API_KEY": "",
"PAPER_SEARCH_MCP_ZENODO_ACCESS_TOKEN": "",
"PAPER_SEARCH_MCP_GOOGLE_SCHOLAR_PROXY_URL": "",
"PAPER_SEARCH_MCP_IEEE_API_KEY": ""
}
}
}
}
Method 4 — pip (standard Python install)
pip install paper-search-mcp
Claude Desktop config:
{
"mcpServers": {
"paper-search-mcp": {
"command": "python",
"args": ["-m", "paper_search_mcp.server"],
"env": {
"PAPER_SEARCH_MCP_UNPAYWALL_EMAIL": "your@email.com",
"PAPER_SEARCH_MCP_CORE_API_KEY": "",
"PAPER_SEARCH_MCP_SEMANTIC_SCHOLAR_API_KEY": "",
"PAPER_SEARCH_MCP_ZENODO_ACCESS_TOKEN": "",
"PAPER_SEARCH_MCP_GOOGLE_SCHOLAR_PROXY_URL": "",
"PAPER_SEARCH_MCP_IEEE_API_KEY": ""
}
}
}
}
If
pythonis not on your PATH, replace it with the full path (e.g./usr/bin/python3orC:\Python311\python.exe). Runwhich python3/where pythonto find it.
Method 5 — npx (via Smithery CLI, no local Python needed)
npx -y @smithery/cli run @openags/paper-search-mcp
Claude Desktop config:
{
"mcpServers": {
"paper-search-mcp": {
"command": "npx",
"args": ["-y", "@smithery/cli", "run", "@openags/paper-search-mcp"],
"env": {
"PAPER_SEARCH_MCP_UNPAYWALL_EMAIL": "your@email.com",
"PAPER_SEARCH_MCP_CORE_API_KEY": "",
"PAPER_SEARCH_MCP_SEMANTIC_SCHOLAR_API_KEY": ""
}
}
}
}
Method 6 — Docker
docker build -t paper-search-mcp .
docker run --rm -i \
-e PAPER_SEARCH_MCP_UNPAYWALL_EMAIL=your@email.com \
-e PAPER_SEARCH_MCP_CORE_API_KEY=your_core_key \
paper-search-mcp
Claude Desktop config:
{
"mcpServers": {
"paper-search-mcp": {
"command": "docker",
"args": ["run", "--rm", "-i", "paper-search-mcp"],
"env": {
"PAPER_SEARCH_MCP_UNPAYWALL_EMAIL": "your@email.com",
"PAPER_SEARCH_MCP_CORE_API_KEY": "",
"PAPER_SEARCH_MCP_SEMANTIC_SCHOLAR_API_KEY": "",
"PAPER_SEARCH_MCP_ZENODO_ACCESS_TOKEN": "",
"PAPER_SEARCH_MCP_GOOGLE_SCHOLAR_PROXY_URL": "",
"PAPER_SEARCH_MCP_IEEE_API_KEY": ""
}
}
}
}
Method 7 — Clone & run from source (development / recommended for macOS local)
This is the most reliable method on macOS — no wrapper scripts, no realpath issues.
# 1. Install uv (skip if already installed)
curl -LsSf https://astral.sh/uv/install.sh | sh
# 2. Clone repo
git clone https://github.com/openags/paper-search-mcp.git
cd paper-search-mcp
# 3. Verify it runs (uv auto-resolves dependencies, no manual install needed)
uv run -m paper_search_mcp.server
Claude Desktop config (replace the directory path with your actual clone location):
{
"mcpServers": {
"paper-search-mcp": {
"command": "uv",
"args": [
"run",
"--directory", "/path/to/paper-search-mcp",
"-m", "paper_search_mcp.server"
],
"env": {
"PAPER_SEARCH_MCP_UNPAYWALL_EMAIL": "your@email.com",
"PAPER_SEARCH_MCP_CORE_API_KEY": "",
"PAPER_SEARCH_MCP_SEMANTIC_SCHOLAR_API_KEY": "",
"PAPER_SEARCH_MCP_ZENODO_ACCESS_TOKEN": "",
"PAPER_SEARCH_MCP_GOOGLE_SCHOLAR_PROXY_URL": "",
"PAPER_SEARCH_MCP_IEEE_API_KEY": ""
}
}
}
}
For example, if you cloned to /Users/mac/Pengsong/paper-search-mcp:
"args": ["run", "--directory", "/Users/mac/Pengsong/paper-search-mcp", "-m", "paper_search_mcp.server"]
uv runautomatically installs dependencies into an isolated environment on first run — nopip installorvenvneeded.
To run one shared network server instead of one stdio process per client:
paper-search-mcp --transport streamable-http --host 127.0.0.1 --port 8000 --path /mcp
The available transports are stdio, sse, and streamable-http. The default
remains stdio. The same network settings can be supplied with
PAPER_SEARCH_MCP_TRANSPORT, PAPER_SEARCH_MCP_HOST,
PAPER_SEARCH_MCP_PORT, and PAPER_SEARCH_MCP_PATH; command-line options take
precedence. Binding to a non-loopback host such as 0.0.0.0 exposes an
unauthenticated server, so place it behind an authenticated gateway rather than
publishing it directly to the internet.
For active development, optionally install an editable copy:
uv venv && source .venv/bin/activate # Windows: .venv\Scripts\activate
uv pip install -e ".[dev]"
Environment Variables (.env file)
Instead of putting keys directly in the JSON config you can store them in the user config file (auto-loaded on startup):
mkdir -p ~/.config/paper-search-mcp
curl -fsSL https://raw.githubusercontent.com/openags/paper-search-mcp/main/.env.example \
-o ~/.config/paper-search-mcp/.env
$EDITOR ~/.config/paper-search-mcp/.env
PAPER_SEARCH_MCP_UNPAYWALL_EMAIL=your@email.com
PAPER_SEARCH_MCP_CORE_API_KEY=
PAPER_SEARCH_MCP_SEMANTIC_SCHOLAR_API_KEY=
PAPER_SEARCH_MCP_ZENODO_ACCESS_TOKEN=
PAPER_SEARCH_MCP_GOOGLE_SCHOLAR_PROXY_URL=
PAPER_SEARCH_MCP_IEEE_API_KEY=
To use a custom path: export PAPER_SEARCH_MCP_ENV_FILE=/absolute/path/to/.env
Legacy variable names without the
PAPER_SEARCH_MCP_prefix (e.g.CORE_API_KEY,UNPAYWALL_EMAIL) are still supported for backward compatibility.
Contributing
We welcome contributions! Here's how to get started:
Fork the Repository: Click "Fork" on GitHub.
Clone and Set Up:
git clone https://github.com/yourusername/paper-search-mcp.git cd paper-search-mcp uv venv && source .venv/bin/activate uv pip install -e ".[dev]"Make Changes:
- Add new platforms in
academic_platforms/. - Update tests in
tests/.
- Add new platforms in
Submit a Pull Request: Push changes and create a PR on GitHub.
Demo

TODO
Planned Academic Platforms
- [√] arXiv
- [√] PubMed
- [√] bioRxiv
- [√] medRxiv
- [√] Google Scholar
- [√] IACR ePrint Archive
- [√] Semantic Scholar
- [√] Crossref
- [√] PubMed Central (PMC)
- [√] CORE
- [√] Europe PMC
- [√] Sci-Hub warning and enablement docs
Development Tasks
- [√] Fix Async search bugs and ensure reliable fast MCP events
- [√] End-to-End full pipeline testing script (search, parse, download)
- [√] Establish two-layer federated architecture (Layer 1 tool:
search_papers) - [√] Ensure pervasive DOI extraction across metadata fields & abstract fallbacks
- Citation graph & Paper relation context feature
- [√] Expand full-stack OpenAlex provider
Priority Free and Open Sources
- [√] PubMed Central (PMC)
- [√] CORE
- [√] OpenAlex
- [√] Europe PMC
- [√] OpenAIRE
- [√] dblp
- [√] CiteSeerX
- [√] DOAJ
- [√] BASE
- [√] Zenodo
- [√] HAL
- [√] SSRN (discovery + best-effort full-text)
- [√] Unpaywall (standalone DOI search source)
Optional and Non-Core Integrations
- ResearchGate
- JSTOR
- ScienceDirect
- Springer Link
- [√] IEEE Xplore (optional skeleton — activate with
IEEE_API_KEY) - [√] ACM Digital Library (keyless Crossref search; publisher PDF access varies)
- Web of Science
- Scopus
Star History
License
This project is licensed under the MIT License. See the LICENSE file for details.
Happy researching with paper-search-mcp! If you encounter issues, open a GitHub issue.