Agent Skills

llm-router

Universal LLM router for AI coding tools. Works with Claude Code, Cursor, Codex, Gemini CLI, Copilot and more. Free-first fallback chain cuts costs 35-80%.

Install

uvx llm-routing
README.md

llm-router routes AI coding prompts across free, budget, and premium model tiers.

llm-router

Stop spending your Claude Pro/Max quota on questions a free model can answer.

llm-router hooks into your coding tool's own lifecycle, reads each prompt before the model does, and drafts an answer on a free or local model first. No API keys needed on a Claude subscription โ€” routing runs through MCP tools and local models.

One caveat: routed answers are advisory by default โ€” Claude still takes the turn unless LLM_ROUTER_ZERO_CLAUDE=1. By default some tool calls wait until a prompt is routed; LLM_ROUTER_ENFORCE=off turns that off. Draft-rate and acceptance figures live in docs/MEASUREMENT.md, with their n, window and conditions.

PyPI PyPI Downloads Tests Stars Listed on RouterArena License

๐Ÿ“‘ Table of Contents

Why install this

You are on a Claude Pro or Max plan, have not spent a cent beyond it, and still hit the five-hour limit โ€” not because you asked too much, but because every prompt went to the premium model: "what does this error mean", "reformat this JSON", "is the service up" drew the same quota as the question that needed it.

Why a proxy cannot do this

Most routers here are a proxy: your agent forwards requests using your API keys. That has a hard limit โ€” a proxy cannot intercept a session authenticated by a subscription, because there is no key to forward.

Pays per token Pays a subscription
What runs out your invoice your five-hour window
Needs API keys yes no
A proxy can help yes no โ€” nothing to intercept
llm-router helps yes yes

Honesty as a feature

llm-router status reports verified and unverified savings separately, each with its own n โ€” not a blended percentage. This project does not publish a general savings percentage; any number depends on your own workload. See Savings below.

Animated benefits panel for llm-router showing cheaper routing, preserved quality, quota protection, and low-config setup.


Quick Start

pip install llm-routing        # installs the `llm-router` command
llm-router install             # wire up Claude Code (default host)
llm-router doctor               # check provider connectivity and setup
llm-router status               # verify it's routing โ€” est. savings vs baseline

Works with no API keys on a Claude Pro/Max subscription โ€” routing goes through MCP tools and local models. Adding provider keys widens the pool; nothing requires them. See guide/PROVIDERS.md. To install into another host, see Works With.


Works With

Host Install
Claude Code (default) llm-router install
Codex CLI llm-router install --host codex
OpenCode llm-router install --host opencode
Gemini CLI llm-router install --host gemini-cli
GitHub Copilot CLI llm-router install --host copilot-cli
OpenClaw llm-router install --host openclaw
Trae IDE llm-router install --host trae
Pi (pi.dev) llm-router install --host pi
Factory Droid llm-router install --host factory
Claude Desktop llm-router install --host desktop
VS Code (native MCP) llm-router install --host vscode
Cursor llm-router install --host cursor
GitHub Copilot in VS Code (capability extension, no cost-routing) llm-router install --host copilot
Windsurf / Cascade llm-router install --host windsurf
Kimi Code (Moonshot AI) llm-router install --host kimi

llm-router install --host all installs or prints every host config in one pass. Full per-host detail, including what each host genuinely cannot do: guide/HOST_SUPPORT_MATRIX.md.


How It Works

Hooks intercept the prompt before your coding tool's own model sees it. A free regex heuristic classifies it instantly, scoring about half of real prompts with confidence โ€” 788 of 1,571 measured (scripts/measure_low_signal_rate.py, run 2026-09-23); the rest fall back to a default. A routable prompt gets a draft from a free/local model first, then walks up the chain toward paid models only if needed.

The draft is advisory โ€” handed to Claude as an unverified hint, not a turn replacement, unless LLM_ROUTER_ZERO_CLAUDE=1 is set. Enforcement acts at the tool-call level: smart (default) holds selected tool calls until the prompt is routed; off disables that. Full modes and per-host overrides: guide/GETTING_STARTED.md.


Features

  • Secrets never leave your machine. A prompt containing an API key, token or private key routes to local models only โ€” fail-closed.
  • Automatic fallback with circuit breakers. A provider that fails or rate-limits is skipped, not retried into the ground.
  • You can see it working. A status line, terminal title and OS notification show the last model routed, savings and health.
  • Session-end summary. Savings vs baseline, tier mix, per-provider cost, latency p50/p95/p99 and top routes.
  • Media and pipelines too. llm_image / llm_video / llm_audio, and llm_orchestrate for multi-step research.

Reference

The default MCP surface shows 12 front-door tools; LLM_ROUTER_SLIM=off shows everything registered.

Topic What Guide
CLI install, status, gain, doctor, okf index/status, sessions status and more guide/GETTING_STARTED.md
Providers Free-first โ€” Ollama (local), OpenRouter, Gemini, Groq, your Claude subscription guide/PROVIDERS.md
Routing Policies conservative โ†’ balanced (default) โ†’ cost_aggressive; thresholds and YAML schema guide/POLICIES.md
MCP Tools Every tool with its signature guide/TOOLS.md
llm-router gain today           # savings so far today
llm-router okf status           # knowledge base freshness
llm-router sessions status      # per-session summary

Savings

Savings figures are a counterfactual โ€” what the same tokens would have cost at API list price, against what was actually spent โ€” not money saved on a flat subscription. llm-router status and llm-router savings-report split verified from unverified savings, each with its own n. Methodology, assumptions and limitations: docs/MEASUREMENT.md.


Trust and Security

llm-router runs entirely on your machine โ€” no hosted proxy, nothing sent to an llm-router service, no account required. A grounding check (src/llm_router/grounding.py) discards any draft citing a file or function absent from its context and the indexed repo, falling through to Claude.

LLM_ROUTER_DIRECT_EXECUTION is default on โ€” a local model gets write_file, edit_file and run_command under a confined, allowlisted path, and none of that stops targeted deletes, git push --force, or exfiltration. Turn it off with LLM_ROUTER_DIRECT_EXECUTION=false. Full analysis: SECURITY.md.

Self-audits are published in audit/ and docs/repo_goals/AUDIT-2026-09-25.md (13 met, 6 partial, 1 not met), including negative results.

Ground Truth accumulation (opt-in, LLM_ROUTER_GROUND_TRUTH=1) builds a corpus of replayable routing decisions instead of unlinked telemetry; off by default since it's the only part that writes prompt text to disk. Details: guide/GROUND_TRUTH.md.


Documentation

Full index: guide/README.md

Document Purpose
Quick Start (2 min) Fastest path to working routing
Getting Started Full setup walkthrough
Host Support Matrix Per-host feature comparison
Providers Provider setup and model recommendations
Routing Policies routing.yaml schema and authoring your own policy
Tool Reference All MCP tools with examples
Architecture Internal design and module structure
Troubleshooting Common issues and fixes
Testing the Router Isolation suite for verifying routing health
Measurement What "savings" means and how it's computed
RouterArena Benchmark methodology, results and negative findings
Ground Truth Replayable routing corpus, opt-in
Benchmarks Model cost/latency/quality table, regenerated by CI
Changelog Release notes (archive)

llm-router is scored on the RouterArena benchmark โ€” full split, 8,400 queries, graded locally with this repo's harness, not yet independently verified. Current Arena Score: 72.35. Full methodology and negative results (including a proxy split that misled tuning by 4.25 points): docs/ROUTERARENA.md.


Enterprise, Contributing, License

llm-router is built for individual developers and small teams: local cost savings, zero ops overhead, no hosted anything. For team-wide policy enforcement, audit export, SSO or per-org budgets, see Chuzom โ€” a separate, sibling product.

Contributions welcome โ€” see CONTRIBUTING.md.

git clone https://github.com/ypollak2/llm-router.git
cd llm-router
uv sync --extra dev
uv run pytest tests/ -q         # Run tests (10,000+)
uv run ruff check src/ tests/   # Lint

MIT License. Issues ยท Discussions ยท PyPI ยท Changelog

Search skills and MCP servers

Fuzzy search across 23,137 skills and servers