Universal LLM router for AI coding tools. Works with Claude Code, Cursor, Codex, Gemini CLI, Copilot and more. Free-first fallback chain cuts costs 35-80%.
Install
uvx llm-routing
llm-router
Stop spending your Claude Pro/Max quota on questions a free model can answer.
llm-router hooks into your coding tool's own lifecycle, reads each prompt before the model does, and drafts an answer on a free or local model first. No API keys needed on a Claude subscription โ routing runs through MCP tools and local models.
One caveat: routed answers are advisory by default โ Claude still takes the
turn unless LLM_ROUTER_ZERO_CLAUDE=1. By default some tool calls wait until a
prompt is routed; LLM_ROUTER_ENFORCE=off turns that off. Draft-rate and
acceptance figures live in docs/MEASUREMENT.md, with
their n, window and conditions.
- Why install this
- Quick Start
- Works With
- How It Works
- Features
- Reference
- Savings
- Trust and Security
- Documentation
- Enterprise, Contributing, License
Why install this
You are on a Claude Pro or Max plan, have not spent a cent beyond it, and still hit the five-hour limit โ not because you asked too much, but because every prompt went to the premium model: "what does this error mean", "reformat this JSON", "is the service up" drew the same quota as the question that needed it.
Why a proxy cannot do this
Most routers here are a proxy: your agent forwards requests using your API keys. That has a hard limit โ a proxy cannot intercept a session authenticated by a subscription, because there is no key to forward.
| Pays per token | Pays a subscription | |
|---|---|---|
| What runs out | your invoice | your five-hour window |
| Needs API keys | yes | no |
| A proxy can help | yes | no โ nothing to intercept |
| llm-router helps | yes | yes |
Honesty as a feature
llm-router status reports verified and unverified savings
separately, each with its own n โ not a blended percentage. This project does
not publish a general savings percentage; any number depends on your own
workload. See Savings below.
Quick Start
pip install llm-routing # installs the `llm-router` command
llm-router install # wire up Claude Code (default host)
llm-router doctor # check provider connectivity and setup
llm-router status # verify it's routing โ est. savings vs baseline
Works with no API keys on a Claude Pro/Max subscription โ routing goes through MCP tools and local models. Adding provider keys widens the pool; nothing requires them. See guide/PROVIDERS.md. To install into another host, see Works With.
Works With
| Host | Install |
|---|---|
| Claude Code (default) | llm-router install |
| Codex CLI | llm-router install --host codex |
| OpenCode | llm-router install --host opencode |
| Gemini CLI | llm-router install --host gemini-cli |
| GitHub Copilot CLI | llm-router install --host copilot-cli |
| OpenClaw | llm-router install --host openclaw |
| Trae IDE | llm-router install --host trae |
| Pi (pi.dev) | llm-router install --host pi |
| Factory Droid | llm-router install --host factory |
| Claude Desktop | llm-router install --host desktop |
| VS Code (native MCP) | llm-router install --host vscode |
| Cursor | llm-router install --host cursor |
| GitHub Copilot in VS Code (capability extension, no cost-routing) | llm-router install --host copilot |
| Windsurf / Cascade | llm-router install --host windsurf |
| Kimi Code (Moonshot AI) | llm-router install --host kimi |
llm-router install --host all installs or prints every host config in one
pass. Full per-host detail, including what each host genuinely cannot do:
guide/HOST_SUPPORT_MATRIX.md.
How It Works
Hooks intercept the prompt before your coding tool's own model sees it. A free
regex heuristic classifies it instantly, scoring about half of real prompts
with confidence โ 788 of 1,571 measured (scripts/measure_low_signal_rate.py,
run 2026-09-23); the rest fall back to a default. A routable prompt gets a
draft from a free/local model first, then walks up the chain toward paid
models only if needed.
The draft is advisory โ handed to Claude as an unverified hint, not a turn
replacement, unless LLM_ROUTER_ZERO_CLAUDE=1 is set. Enforcement acts at the
tool-call level: smart (default) holds selected tool calls until the prompt
is routed; off disables that. Full modes and per-host overrides:
guide/GETTING_STARTED.md.
Features
- Secrets never leave your machine. A prompt containing an API key, token or private key routes to local models only โ fail-closed.
- Automatic fallback with circuit breakers. A provider that fails or rate-limits is skipped, not retried into the ground.
- You can see it working. A status line, terminal title and OS notification show the last model routed, savings and health.
- Session-end summary. Savings vs baseline, tier mix, per-provider cost, latency p50/p95/p99 and top routes.
- Media and pipelines too.
llm_image/llm_video/llm_audio, andllm_orchestratefor multi-step research.
Reference
The default MCP surface shows 12 front-door tools; LLM_ROUTER_SLIM=off
shows everything registered.
| Topic | What | Guide |
|---|---|---|
| CLI | install, status, gain, doctor, okf index/status, sessions status and more |
guide/GETTING_STARTED.md |
| Providers | Free-first โ Ollama (local), OpenRouter, Gemini, Groq, your Claude subscription | guide/PROVIDERS.md |
| Routing Policies | conservative โ balanced (default) โ cost_aggressive; thresholds and YAML schema |
guide/POLICIES.md |
| MCP Tools | Every tool with its signature | guide/TOOLS.md |
llm-router gain today # savings so far today
llm-router okf status # knowledge base freshness
llm-router sessions status # per-session summary
Savings
Savings figures are a counterfactual โ what the same tokens would have cost
at API list price, against what was actually spent โ not money saved on a flat
subscription. llm-router status and llm-router savings-report split
verified from unverified savings, each with its own n. Methodology,
assumptions and limitations: docs/MEASUREMENT.md.
Trust and Security
llm-router runs entirely on your machine โ no hosted proxy, nothing sent to an
llm-router service, no account required. A grounding check
(src/llm_router/grounding.py) discards any draft citing a file or function
absent from its context and the indexed repo, falling through to Claude.
LLM_ROUTER_DIRECT_EXECUTION is default on โ a local model gets write_file,
edit_file and run_command under a confined, allowlisted path, and none of
that stops targeted deletes, git push --force, or exfiltration. Turn it off
with LLM_ROUTER_DIRECT_EXECUTION=false. Full analysis:
SECURITY.md.
Self-audits are published in audit/ and docs/repo_goals/AUDIT-2026-09-25.md (13 met, 6 partial, 1 not met), including negative results.
Ground Truth accumulation (opt-in, LLM_ROUTER_GROUND_TRUTH=1) builds a
corpus of replayable routing decisions instead of unlinked telemetry; off by
default since it's the only part that writes prompt text to disk. Details:
guide/GROUND_TRUTH.md.
Documentation
Full index: guide/README.md
| Document | Purpose |
|---|---|
| Quick Start (2 min) | Fastest path to working routing |
| Getting Started | Full setup walkthrough |
| Host Support Matrix | Per-host feature comparison |
| Providers | Provider setup and model recommendations |
| Routing Policies | routing.yaml schema and authoring your own policy |
| Tool Reference | All MCP tools with examples |
| Architecture | Internal design and module structure |
| Troubleshooting | Common issues and fixes |
| Testing the Router | Isolation suite for verifying routing health |
| Measurement | What "savings" means and how it's computed |
| RouterArena | Benchmark methodology, results and negative findings |
| Ground Truth | Replayable routing corpus, opt-in |
| Benchmarks | Model cost/latency/quality table, regenerated by CI |
| Changelog | Release notes (archive) |
llm-router is scored on the RouterArena
benchmark โ full split, 8,400 queries, graded locally with this repo's harness,
not yet independently verified. Current Arena Score: 72.35. Full
methodology and negative results (including a proxy split that misled tuning
by 4.25 points): docs/ROUTERARENA.md.
Enterprise, Contributing, License
llm-router is built for individual developers and small teams: local cost
savings, zero ops overhead, no hosted anything. For team-wide policy
enforcement, audit export, SSO or per-org budgets, see
Chuzom โ a separate, sibling product.
Contributions welcome โ see CONTRIBUTING.md.
git clone https://github.com/ypollak2/llm-router.git
cd llm-router
uv sync --extra dev
uv run pytest tests/ -q # Run tests (10,000+)
uv run ruff check src/ tests/ # Lint
MIT License. Issues ยท Discussions ยท PyPI ยท Changelog
