Agent Skills

mineru-mcp

docslinxule10 stars

MinerU document parsing API — PDFs, images, DOCX, PPTX with OCR and batch processing.

Install

npx -y mineru-mcp
README.md

mineru-mcp

MCP server and mineru-cloud CLI for the MinerU document parsing API, with complete artifact retention, durable operations, and portable PDF bundles.

Durable operations and bundles in this checkout are unreleased. The published mineru-mcp@1.2.0 does not include them; npm users need a release containing these changes.

Features

  • VLM model — Document parsing with a vision-language model
  • Pipeline model — Document parsing with the pipeline model
  • Local file upload — Upload files from disk for batch parsing
  • Batch processing — Submit multiple documents in one request
  • Download & rename — Extract markdown with original filenames
  • Page ranges — Extract specific pages only
  • Long documents — mineru_parse_long submits page-range slices and mineru_merge_slices stitches
  • CLI twin — mineru-cloud calls the same tools from a shell
  • Language and OCR options — Select a language and enable OCR for the pipeline model

Tools

Tool Description
mineru_parse Parse a document URL
mineru_status Check task progress, get download URL
mineru_batch Parse multiple URLs
mineru_batch_status Get batch results with pagination
mineru_upload_batch Upload local files for batch parsing
mineru_download_results Retain complete archives, inventory all members, and create named compatibility copies
mineru_parse_long Submit a long document as a batch of page-range slices
mineru_merge_slices Join available Markdown and retain immutable slice archives with explicit unknown page provenance
mineru_capabilities Read locally known capabilities; explicit refresh performs discovery
mineru_submit Submit with a durable local journal and duplicate-request protection
mineru_operation_status Read a recorded operation; explicit refresh checks remote status
mineru_resume Continue safe recorded checkpoints
mineru_cancel Stop local processing; remote cancellation is unsupported
mineru_bundle Validate and return an operation's retained bundle

Installation

Use Node.js 24 for new installations. Node.js 18+ remains supported for compatibility. Cloud commands require a MinerU API key; offline bundle creation requires no credentials.

Local output retention, bundles, and durable operations require macOS or Linux. These paths use directory-bound filesystem operations and fail closed on Windows; the Windows client configuration below does not imply local retention support.

CLI Install (one-liner)

# Claude Code
claude mcp add mineru-mcp -e MINERU_API_KEY=your-api-key -- npx -y mineru-mcp

# Codex CLI (OpenAI)
codex mcp add mineru --env MINERU_API_KEY=your-api-key -- npx -y mineru-mcp

# Gemini CLI (Google)
gemini mcp add -e MINERU_API_KEY=your-api-key mineru npx -y mineru-mcp

Claude Desktop

Add to your claude_desktop_config.json:

OS Config path
macOS ~/Library/Application Support/Claude/claude_desktop_config.json
Windows %APPDATA%\Claude\claude_desktop_config.json
Linux ~/.config/Claude/claude_desktop_config.json
{
  "mcpServers": {
    "mineru": {
      "command": "npx",
      "args": ["-y", "mineru-mcp"],
      "env": {
        "MINERU_API_KEY": "your-api-key"
      }
    }
  }
}

VS Code

Add to .vscode/mcp.json (workspace) or open Command Palette > MCP: Open User Configuration (global):

{
  "servers": {
    "mineru": {
      "command": "npx",
      "args": ["-y", "mineru-mcp"],
      "env": {
        "MINERU_API_KEY": "your-api-key"
      }
    }
  }
}

Note: VS Code uses "servers" as the top-level key, not "mcpServers". Other VS Code forks (Trae, Void, PearAI, etc.) typically use this same format.

Cursor

Add to ~/.cursor/mcp.json (global) or .cursor/mcp.json (project):

{
  "mcpServers": {
    "mineru": {
      "command": "npx",
      "args": ["-y", "mineru-mcp"],
      "env": {
        "MINERU_API_KEY": "your-api-key"
      }
    }
  }
}

Windsurf

Add to ~/.codeium/windsurf/mcp_config.json (Windows: %USERPROFILE%\.codeium\windsurf\mcp_config.json):

{
  "mcpServers": {
    "mineru": {
      "command": "npx",
      "args": ["-y", "mineru-mcp"],
      "env": {
        "MINERU_API_KEY": "your-api-key"
      }
    }
  }
}

Cline

Open MCP Servers icon in Cline panel > Configure > Advanced MCP Settings, then add:

{
  "mcpServers": {
    "mineru": {
      "command": "npx",
      "args": ["-y", "mineru-mcp"],
      "env": {
        "MINERU_API_KEY": "your-api-key"
      }
    }
  }
}

Cherry Studio

In Settings > MCP Servers > Add Server, set Type to STDIO, Command to npx, Args to -y mineru-mcp, and add environment variable MINERU_API_KEY. Or paste in JSON/Code mode:

{
  "mineru": {
    "name": "MinerU",
    "command": "npx",
    "args": ["-y", "mineru-mcp"],
    "env": {
      "MINERU_API_KEY": "your-api-key"
    },
    "isActive": true
  }
}

Witsy

In Settings > MCP Servers, add a new server with Type: stdio, Command: npx, Args: -y mineru-mcp, and set environment variable MINERU_API_KEY to your API key.

Codex CLI (TOML config)

Alternatively, edit ~/.codex/config.toml directly:

[mcp_servers.mineru]
command = "npx"
args = ["-y", "mineru-mcp"]

[mcp_servers.mineru.env]
MINERU_API_KEY = "your-api-key"

Gemini CLI (JSON config)

Alternatively, edit ~/.gemini/settings.json directly:

{
  "mcpServers": {
    "mineru": {
      "command": "npx",
      "args": ["-y", "mineru-mcp"],
      "env": {
        "MINERU_API_KEY": "your-api-key"
      }
    }
  }
}

Windows

On Windows, npx requires a shell wrapper. Replace "command": "npx" with:

{
  "command": "cmd",
  "args": ["/c", "npx", "-y", "mineru-mcp"],
  "env": {
    "MINERU_API_KEY": "your-api-key"
  }
}

For CLI tools on Windows:

claude mcp add mineru-mcp -e MINERU_API_KEY=your-api-key -- cmd /c npx -y mineru-mcp
codex mcp add mineru --env MINERU_API_KEY=your-api-key -- cmd /c npx -y mineru-mcp

ChatGPT

ChatGPT only supports remote MCP servers over HTTPS — local stdio servers like this one are not directly supported. You would need to deploy behind a public URL with HTTP transport.

CLI: mineru-cloud

Every tool is also a shell command — the CLI runs the MCP server in-process over an in-memory transport, so the two can't drift. Same env vars (MINERU_API_KEY, MINERU_BASE_URL, MINERU_DEFAULT_MODEL).

mineru-cloud list                                    # commands + options (from the tool schemas)
mineru-cloud parse --url https://arxiv.org/pdf/2303.08774 --pages 1-10
mineru-cloud status --task-id <id> --wait            # --wait polls every 10s until done/failed
mineru-cloud batch --urls '["https://…/a.pdf","https://…/b.pdf"]'
mineru-cloud download-results --batch-id <id> --output-dir ./papers --wait

# Long document: slice, then stitch
mineru-cloud parse-long --url https://…/book.pdf --total-pages 520 --name book
mineru-cloud merge-slices --batch-id <id> --output-dir ./books --wait

Options mirror the tool parameters with _ → - (--total-pages, --output-dir); numbers, true/false and JSON arrays are coerced. Install: bun add -g mineru-mcp (or npm i -g).

Offline artifact bundles

Create a portable Scholia v1.0.1 bundle from an existing provider ZIP and the exact source PDF. This command runs offline and does not require an API key:

mineru-cloud bundle --source /absolute/source.pdf --archive /absolute/result.zip \
  --output /absolute/bundles --json

The command returns bundle_dir, an immutable directory containing bundle.json, the original PDF, and the byte-identical ZIP. Move the entire directory for import. Repeated inputs verify the existing bytes and return the same bundle. Optional --batch-id and --model record caller-supplied metadata. --binding caller_asserted records a qualified association; the default is unknown. Neither choice proves that the provider parsed those exact PDF bytes. The command does not infer page coverage, requested options, or provider versions from filenames.

Only PDF sources qualify for this bundle. Existing non-PDF cloud commands remain available and produce diagnostic inventories. Unsafe archives are retained in quarantine, outside importable bundles. ZIP64, encrypted archives, special files, ambiguous legacy filename encodings, and unsafe or colliding paths are rejected.

Complete downloads and structured status

download-results retains each ZIP under {name}/archives/<sha256>.zip and writes an inventory.json covering every member, including unknown formats. It also creates legacy {name}.md, {name}_content.json, and image copies when selection is unambiguous. A missing or ambiguous Markdown file does not discard the archive. No shell unzip executable is used. Inventory hashes expanded bytes incrementally; selected compatibility copies have a separate 64 MiB size limit.

Use --json on lifecycle commands for structured operation IDs, normalized state, pollability, counts, and errors. --wait uses typed state across the whole batch, including entries outside the displayed page. It polls every 10 seconds for up to 30 minutes. A failed entry does not hide other pending entries. Unknown provider states remain explicit. Exit status is 1 for failure and 2 for partial or unknown results. The 8 existing MCP tool names and their CLI commands remain available.

Slice merges retain each archive under a hash-qualified directory. Earlier image links remain valid after a successor merge. The merged content JSON is a slice receipt with archive references, rather than a flattened array with assumed page offsets. Missing Markdown or failed slices produce an explicit partial result; otherwise coverage and original PDF page provenance remain unknown.

Durable standalone operations

The 6 additional commands share their implementation with MCP:

mineru-cloud capabilities --api v4 --json
mineru-cloud submit --file /absolute/source.pdf --api v4 --model vlm --output-dir /absolute/results --json
mineru-cloud operation-status --operation-id ID --json
mineru-cloud resume --operation-id ID --wait --wait-timeout-seconds 1800 --json
mineru-cloud cancel --operation-id ID --json
mineru-cloud bundle --operation-id ID --json

MINERU_STATE_DIR selects the journal root; the default is ~/.local/share/mineru-cloud. Keep that directory for recovery. The journal retains source bytes, request fingerprints, remote IDs, completed output hashes, and phase checkpoints. Credentials and signed URLs stay in process memory. Exact duplicate submissions join the recorded operation. A lost allocation or job ID requires reconciliation and never triggers automatic resubmission. An ambiguous V4 upload with a saved batch ID resumes by polling that batch. V1 inspects a known upload before continuing an uncertain completion.

One process holds the writer lock at a time. Normal dead-process locks can be reclaimed with an owner-token check. An interrupted lock-recovery step or a reused PID fails closed with an explicit busy/recovery error. This filesystem journal is separate from Scholia's SQLite lease and reservation system; it does not claim Scholia's worker scheduling or maintenance guarantees.

operation-status reads locally unless --refresh is supplied. resume continues a safe checkpoint. Finalization uses saved outputs without contacting the provider. Local cancellation does not cancel remote processing or remove uploaded data. Remote cancellation and automatic lost-ID association remain unsupported. Completed bundles preserve archives and every detached output, including unknown formats. Provider transport completion does not establish full PDF page coverage.

When one output remains unavailable after bounded download attempts, the operation publishes a partial bundle containing the verified successes and typed failure details. bundle --operation-id ID can return that retained evidence. An explicit later resume may recover missing outputs from the same operation and publish a successor; earlier bundles remain immutable. A partial result is not permission to submit the document again.

New operation commands use a structured Result envelope in CLI JSON and MCP structuredContent, while retaining the existing flattened fields for compatibility. Their CLI exit codes are 0 for ok, 1 for partial or error, and 2 for invalid arguments. The original eight commands retain their historical error=1 and partial=2 exit codes. Offline bundle --source ... --archive ... also retains its existing created/existing receipt shape. Check status, state, and recovery details rather than interpreting a nonzero exit as permission to resubmit.

The process-restart tests kill workers at actual source, submission-intent, output-retention, and finalization checkpoints. Verified source orphans created before the first journal can be adopted safely; uncertain remote submissions cannot. Source/output writes are synced before their journal references. These tests establish bounded process-crash behavior, not a universal power-loss or network-filesystem durability guarantee.

The modern V1 adapter uses uploads, parse/jobs, and files/{id}/content. It does not use the older agent/parse API. Explicit capability refresh queries health and the separate tiers endpoint. No V4 model is mapped to a V1 tier. V1 range parsing is unsupported. Hosted V1 execution remains disabled by default until endpoint-specific validation; fixture tests can inject an adapter.

New network calls require public HTTPS addresses. DNS results are checked and pinned at connection time; redirects are revalidated and lose API authorization. Transfer URLs receive only explicitly supplied headers. Private/self-hosted endpoints require a separately implemented trust policy and are not enabled here. New submit currently accepts PDFs; the original 8 commands retain non-PDF support. There is no account quota reservation, automatic scheduler, or operator UI for uncertain-operation association in this standalone milestone.

Official protocol references are the pinned V1 guide, pinned example, and pinned API server schema. The current guide was checked on September 30, 2026. All new execution behavior is fixture-tested; no live provider validation, version bump, or release is claimed.

Configuration

Environment Variable Default Description
MINERU_API_KEY (required) Your MinerU API Bearer token
MINERU_BASE_URL https://mineru.net/api/v4 API base URL
MINERU_DEFAULT_MODEL pipeline Default model: pipeline or vlm

Get your API key at mineru.net

Usage

Parse a single URL

mineru_parse({
  url: "https://example.com/document.pdf",
  model: "vlm",        // optional: "pipeline" (default) or "vlm"
  pages: "1-10,15",    // optional: page ranges
  ocr: true,           // optional: enable OCR (pipeline only)
  formula: true,       // optional: formula recognition
  table: true,         // optional: table recognition
  language: "en",      // optional: language code
  formats: ["html"]    // optional: extra export formats
})

Check task progress

mineru_status({
  task_id: "abc-123",
  format: "concise"    // optional: "concise" (default) or "detailed"
})

Concise output: done | abc-123 | https://cdn-mineru.../result.zip

Batch parse URLs

mineru_batch({
  urls: ["https://example.com/doc1.pdf", "https://example.com/doc2.pdf"],
  model: "vlm"
})

Check batch progress

mineru_batch_status({
  batch_id: "batch-123",
  limit: 10,           // optional: max results (default: 10)
  offset: 0,           // optional: skip first N results
  format: "concise"    // optional: "concise" or "detailed"
})

Upload local files

mineru_upload_batch({
  directory: "/path/to/pdfs",  // scan directory for supported files
  // OR
  files: ["/path/to/doc1.pdf", "/path/to/doc2.pdf"],  // explicit file list
  model: "vlm",        // optional
  formula: true,       // optional
  table: true,         // optional
  language: "en",      // optional
  formats: ["html"]    // optional
})

Returns batch_id for tracking. Each file's original name is preserved via data_id (spaces become underscores).

Download results as markdown

mineru_download_results({
  batch_id: "batch-123",       // from mineru_upload_batch or mineru_batch
  output_dir: "/path/to/output",
  overwrite: false             // optional: overwrite existing files
})

Output filenames are derived from data_id (e.g., my_paper_title.md). Spaces in original filenames become underscores.

Typical local file workflow

mineru_upload_batch → mineru_batch_status (poll) → mineru_download_results

Supported Formats

  • PDF, DOC, DOCX, PPT, PPTX
  • PNG, JPG, JPEG

Local limits and provider evidence

Legacy batch commands enforce at most 200 items; legacy local uploads reject files over 200 MiB. parse-long limits each requested slice to 200 pages. These are current implementation guards, not verification of a provider account's current limits, quota, queue priority, or supported plan. Check those separately when a cloud task is authorized. This checkout's synthetic tests do not measure hosted parsing accuracy, speed, or language coverage.

Release 1.1.6

Restores Node.js 18 HTTP compatibility for fresh installs by retaining MCP SDK 1.29.x and its Node 18-compatible Hono adapter. SDK 1.30 permits an adapter that requires Node.js 20. Version 1.1.5 passed the locked dependency checks but the published-package check exposed an HTTP initialization failure on a fresh install. CI now installs the packed package without the repository lock and exercises both transports on Node.js 18. The SDK compatibility bound is intentional; revisit it with this consumer-install gate before adopting a newer SDK.

Release 1.1.5

Maintenance release: audited dependency updates, Express 5 and Zod 4 compatibility, and regression coverage for both transports. The MCP handshake and HTTP startup message now report the package version instead of the stale 1.0.2 value. Tool inputs and document-processing behavior are unchanged.

Development

Read AGENTS.md for the canonical contributor handoff, task authority, recovery invariants and paired Scholia fixture/integration checks.

Use Bun 1.4.2 and the recommended Node.js 24 runtime for development:

bun install --frozen-lockfile
bun audit
bun run build
bun run test
bun run test:package

The tests use synthetic fixtures and local MinerU API doubles. They exercise the built stdio and HTTP servers, durable operation recovery, exact artifact retention, bundle identity, and filesystem containment without real credentials or provider calls. Temporary fixtures use the canonical system temp directory on macOS and Linux. These checks do not establish live-provider parsing accuracy.

On PRs to main, version-tag pushes and manual registry recovery, CI audits dependencies, builds, and runs the full suite on both macOS and Linux with Node.js 24. Each platform also installs a fresh packed package without the repository lock and tests it on Node.js 18. This consumer-install gate preserves the Node 18 compatibility target. Dependabot updates the Bun manifest and lockfile together. Ordinary pushes to main do not trigger this workflow.

Publishing

With explicit release authorization, bump package.json, the top-level version in server.json, and its matching npm package versions, complete the checks above, merge, then push the matching vX.Y.Z tag. Before publication, CI requires the tag and all of these versions to match. The existing ci check requires every macOS/Linux verification job to pass before the separate publishing job can run. CI resolves the requested checkout once, then verifies and publishes that exact commit. Publication uses Node.js 24 and GitHub OIDC: it publishes to npm, waits for the exact package version to become available, then registers it with the MCP Registry. If registry registration fails after npm succeeds, retry only registration using the existing immutable tag:

gh workflow run publish-mcp.yml --ref main -f registry_tag=vX.Y.Z

Replace vX.Y.Z with the already-published tag. Recovery resolves refs/tags/vX.Y.Z, verifies that tag's checkout, skips npm publication, and retries only MCP Registry registration.

License

MIT

Links

Search skills and MCP servers

Fuzzy search across 23,137 skills and servers