CodeNib
Find source context and trace calls across languages in Claude Code, Codex and other MCP agents.
Install
uvx codenibFind the right code. Trace the calls.
Give Claude Code and Codex source context and typed graph navigation across languages.
71.4% code-block Recall@5 on 100 real issues.
Model-planned grep → Jev: +12.8 percentage points over the same grep candidates without reranking.
Experimental result · Method and limitations
Use with your agent · Browse a Wiki · Languages · Reproductions · Releases
▶ Watch the 15-second replay · real CLI output · waits condensed · Pinned source, transcript and setup details.
Quickstart
Requires Python 3.10+, Git, a clean repository, and Claude Code or Codex.
python -m pip install "codenib[graph,mcp]"
codenib codegraph init /path/to/your/repo
Then ask your agent: “Use CodeNib's explore_context to find where request
retry behavior is implemented. Cite the files and lines.”
This shipping path builds local search and a typed symbol graph, then connects installed agent clients. It needs no model, API key, or GPU for CodeNib; your agent uses its own model. It is separate from the grep/Jev experiment above. Language toolchain and project prerequisites vary.
Setup and troubleshooting · Source exclusions · All MCP tools · Local Wiki
Why CodeNib
- Inspect the evidence. Retrieve bounded code context with file and line references, then follow callers, callees, definitions, and references.
- Use the same tools across languages. The registry tracks 14 language entries, including 12 with graph backends. Check capability and setup differences in the language matrix.
- Compare methods on the same source. Pinned agent contracts, datasets, and scorer checks make the reproduction surface inspectable.
- Share what you found. Export a Wiki with source citations to your own GitHub Pages, or browse the public examples.
How it compares
What an agent can ask each tool, checked against upstream docs on 2026-09-28. This compares capabilities, not benchmark scores.
| Tool | Ask in plain language | Callers/callees beyond one hop | Relationships resolved by | Runs |
|---|---|---|---|---|
| grep / read | No; literal or regex | No | Text match | Local |
| CodeNib CodeGraph | Yes; ranked source anchors (explore_context) |
Yes; one call, depth ≤ 8 (dependency_subgraph) |
SCIP/LSP indexers (compiler-resolved) | Local |
| Serena | No; symbol name or regex | One level per call (find_referencing_symbols) |
Live language server or IDE | Local |
| CodeGraph (colbymchenry) | Yes; FTS5 full-text (codegraph_explore) |
Yes; call paths in codegraph_explore |
Tree-sitter AST extraction | Local; telemetry opt-out |
| DeepWiki public MCP | Yes; generated answers (ask_question) |
No structured graph tools | Generated Wiki | Hosted |
None of these repository tools needs its own model or API key; your agent uses its own model. Detailed comparison, sources and boundaries. CodeNib 0.2.4 also includes an optional grep → Jev route using OpenRouter for planning and reranking; selected code goes to remote models. Authorization previews remain opt-in. The historical research result above is separate from the product evaluation and the model-free CodeGraph row.
Languages
Graph backends: Python, Go, Rust, C/C++, C#, Java, Ruby, PHP, Kotlin, Scala, JavaScript, and TypeScript. Swift and Lua support chunking/retrieval. Provider prerequisites and coverage differ by language.
The generated language matrix separates chunking, graph backends, incremental-backend support, and decoder parity. The product currently reuses or rebuilds views; file-level delta repair is not enabled.
Results and reproductions
| Start here | What you can inspect |
|---|---|
| grep → Jev result | 100-issue retrieval comparison, candidate controls, model use and limitations |
| Agent integration matrix | Revision-pinned LocAgent, Agentless, CoSIL, OrcaLoca and RepoNavigator contracts; compatibility does not imply reproduced paper scores |
| Dataset and benchmark matrix | CodeNib Base/Synthesis, SWE-bench variants, Loc-Bench and SWE-Explore support |
| SWE-Explore validation | 1,020/1,020 real-output metric cells match the pinned official evaluator on a fixed 20-case run |
| DGX Spark deployment | Local Wiki, CodeGraph and model serving on GB10; a deployment guide, not a token-saving benchmark |
Documentation and community
Documentation · Architecture · Contributing · CI and testing · Changelog · Discord
CodeNib is in beta; public interfaces may change before a stable release. Historical research artifacts retain their published dataset identifiers.
Citation
If you use CodeNib in your research, please cite our arXiv paper:
@misc{yu2026codenibmultiviewdataserving,
title={CodeNib: A Multi-View Data System for Serving Repository Context to Coding Agents},
author={Zhongming Yu and Hengjia Yu and Boqin Yuan and Shuting Zhao and Yizhao Chen and Aryan Dokania and Mihir Jagtap and Jiayu Chang and Yitong Ma and Yash Jayswal and Wentao Ni and Hejia Zhang and Zhaoling Chen and Gangda Deng and Jishen Zhao},
year={2026},
eprint={2607.25431},
archivePrefix={arXiv},
primaryClass={cs.SE},
url={https://arxiv.org/abs/2607.25431},
}
CodeNib is licensed under Apache 2.0.