Agent Skills

continuum

CONTINUUM: Verifiable semantic recovery for long-running AI agents. Semantic checkpoints (not conversation dumps), an idempotent action ledger that refuses duplicate side effects, and a hash-chained tamper-evident event log, all exposed as a deny-by-default MCP server. Framework-agnostic, Python 3.11+.

Install

uvx continuum-agent
README.md

CONTINUUM Banner

CONTINUUM: Verifiable semantic recovery for long-running AI agents. Semantic checkpoints (not conversation dumps), an idempotent action ledger that refuses duplicate side effects, and a hash-chained tamper-evident event log, all exposed as a deny-by-default MCP server. Framework-agnostic, Python 3.11+.

Python 3.11+ PyPI License Pydantic v2 Website Demo CI status Coverage Stars

Visit the CONTINUUM website

Build with Ona

If CONTINUUM helps your agents recover, please star the repo. It helps others discover it and keeps good first issues coming.

English | 简体中文 | Español | 日本語 | Português | 한국어


Contents

Why · Quick Start · How it works · Where CONTINUUM sits · Features · Security Extension · Empirical Verification · MCP Integration · Framework Integration · Core Concepts · Architecture · API and CLI · Roadmap · What CONTINUUM Is Not · Related work · Status and limitations · Contributing · License


Why

Modern AI agents run long tasks (hundreds of LLM calls, tool invocations, file and database writes). When they crash, the usual response is to replay everything from scratch, which duplicates work, duplicates side effects, wastes tokens, and loses decisions.

CONTINUUM asks a narrower, harder question: can an agent resume from a compact semantic representation of its task state while independently verifying that state is still valid in the current environment? Its differentiator is three-part:

  • Semantic checkpoints: a compact, versioned representation of what the agent needs to continue, not a conversation dump.
  • Independent environment revalidation: every checkpoint component is verified against the current environment before resume, with staleness propagating through the dependency graph. A run can configure the observers it trusts (continuum providers) so resume captures the world by name instead of relying on a caller to remember to pass them; a provider that cannot report marks its resources unknown, never as unchanged.
  • Provenance-aware state: every fact traces to its origin, so agent-reported progress is never self-certifying.

Quick Start

Published to PyPI as continuum-agent 0.1.2: pip install continuum-agent (pip install continuum-agent==0.1.2 to pin). Release tags additionally ship built wheels attached to GitHub Releases.

Paths that need no clone and no local install:

Path How
Install from PyPI pip install continuum-agent==0.1.2, then continuum --help
Install with Homebrew brew install wized2/continuum/continuum, then continuum --help
Watch crash recovery happen end to end docker run --rm ghcr.io/cyrax321/continuum
Use the CLI through Docker docker run --rm ghcr.io/cyrax321/continuum continuum --help
Run the CLI without cloning uvx --from git+https://github.com/Cyrax321/CONTINUUM.git continuum --help
Windows PowerShell (from a clone) powershell -ExecutionPolicy Bypass -File .\try-it.ps1 or powershell -ExecutionPolicy Bypass -File .\try-it.ps1 cli --help
Watch the same recovery in a notebook Open In Colab
The same notebook, on Binder Binder
Full dev environment in the browser Open in GitHub Codespaces

The Docker image is published to GHCR by CI on every push to main and every release tag (.github/workflows/docker-publish.yml). The Codespace is defined in .devcontainer/. The notebook is examples/demo.ipynb: its first cell installs CONTINUUM only when the import fails, so the one file runs on Colab, on Binder, and from a clone.

git clone https://github.com/Cyrax321/CONTINUUM.git
cd CONTINUUM

uv venv && source .venv/bin/activate     # macOS / Linux; Windows: .venv\Scripts\activate

# Contributors (recommended): library + CLI + all test tooling + every adapter
uv pip install -e ".[dev]"

# Or pick only what you need: . (minimal), [mcp], [otel], [langgraph],
# [openai], [langchain], [attest], [postgres]

# Or skip the clone entirely:
uv pip install git+https://github.com/Cyrax321/CONTINUUM.git
uv pip install "continuum-agent[mcp] @ git+https://github.com/Cyrax321/CONTINUUM.git"

pip fallback: replace uv pip install with pip install in every command above.

Verify:

continuum --help                 # CLI entrypoint
continuum-mcp --help             # MCP server entrypoint (needs [mcp] or [dev])
pytest -q                        # ~2,864 collected, ~2,825 passed, ~34 skipped on a minimal env (exact counts vary)
ruff check src/ tests/ examples/ && ruff format --check src/ tests/ examples/
mypy src/continuum               # the three gates CI enforces

The core library has one runtime dependency (pydantic>=2.7); everything else is opt-in. The full package map, extras matrix, Postgres test setup, and per-command verification are in references/install.md.

Wire a coding agent in two minutes

For Claude Code, Gemini CLI, or Codex, you do not write Python and do not need a prompt file:

continuum start my-task --goal "What the agent should do"
continuum hooks install claude-code --with-gate   # also: gemini, codex

From then on every file the agent writes is captured as hash-chained evidence, its session starts with an automatic status briefing, unclaimed side effects registered in .continuum/gate.json are refused before they fire, and a fresh session after any crash resumes with executable next steps. No CLAUDE.md required.

Minimal library example, record and recover:

from continuum import EventType, Run, SQLiteStorage, project

store = SQLiteStorage("agent.db")
store.create_run(Run(run_id="run_4821", goal="Analyze 10,000 documents"))
store.append_event("run_4821", EventType.RUN_STARTED, {"goal": "Analyze 10,000 documents", "total": 10_000})

for i, doc in enumerate(documents):
    analyze(doc)
    store.append_event("run_4821", EventType.WORK_COMPLETED, {"doc": i})

# After a crash, a new process picks up exactly where it stopped:
state = project("run_4821", store.read_events("run_4821"))
print(state.progress.completed)            # already done, not repeated
print(store.verify_events("run_4821").ok)  # True, chain intact after the crash

Run the proof yourself:

python examples/crash_recovery_agent.py   # real process kill, real side effect
python examples/context_compaction.py     # transcript lost, checkpoint survives
python examples/model_switch.py           # Model A dies, Model B resumes safely
python scripts/mcp_smoke.py               # real subprocess, real JSON-RPC traffic

The e2e-autonomy-test/ kit scripts a real invoice-batch task, a hard-kill mid-run, and a fresh resume session, then scores the outbox, ledger, and event chain out of band. Run 1 scored 7/7 mechanics against a real Claude Code session. Full walkthrough in references/e2e.md.

How it works

CONTINUUM separates LLM context (temporary) from durable task state (permanent). Instead of saving conversation history, it constructs a semantic checkpoint, the minimum verified information required to continue.

CONTINUUM how it works

The detailed explanation, the projection model, and the recovery context are in references/architecture.md.

Where CONTINUUM sits

Four concerns overlap in every long-running agent. CONTINUUM owns only the last one and touches the other three through explicit seams. No competitor is named and no claim is made without a shipped module or a published suite that already prints it.

Layer Answers How it connects (shipped modules or published output)
Harness How does the agent call tools and make progress toward a goal? Outside CONTINUUM. Wiring points ship in src/continuum/adapters/generic.py (GenericAgentAdapter), src/continuum/adapters/thin.py (CrewAI, AutoGen, Pydantic AI hooks), src/continuum/mcp/server.py (MCP stdio), src/continuum/hooks.py and src/continuum/clienthooks.py (coding-CLI lifecycle hooks), src/continuum/gateway.py (enforcing HTTP proxy for any language), and src/continuum/otel.py (OpenTelemetry bridge). Recipes are in docs/recipes/ and references/adapters.md.
Durable execution What happened before a crash and what can be replayed without losing work? Hash-chained event log src/continuum/events.py with verify() and trusted_through, durable storage src/continuum/storage/sqlite.py (WAL, synchronous=FULL, schema v6) and src/continuum/storage/postgres.py plus src/continuum/storage/migrations.py, policy-driven checkpoints src/continuum/checkpoint/manager.py and src/continuum/checkpoint/policy.py that replay the gap on restore(). Walkthrough is in docs/recovery_walkthrough.md (output of examples/recovery_walkthrough.py).
Control plane Which run is active, who may act on it, and where does output go? Run registry and parent/child hierarchy src/continuum/storage/ and src/continuum/recovery/family.py (continuum tree), allowlist authz src/continuum/mcp/authz.py (CONTINUUM_MCP_MUTATING_CLIENTS / CONTINUUM_MCP_TOKEN), presentation surfaces src/continuum/dashboard/app.py and src/continuum/serve/server.py, CLI src/continuum/cli/main.py (continuum runs, continuum tree, continuum health).
Verification substrate Given the checkpoint at time T and the world as it is now, is it still safe and correct to continue? src/continuum/state/validator.py (staleness dependency -> evidence -> finding -> decision plus PlanStep.depends_on), src/continuum/provenance_map.py (Origin to REQUIRES_REVIEW until REVIEW_CONFIRMED), src/continuum/actions/ledger.py with src/continuum/actions/idempotency.py and src/continuum/gate.py / src/continuum/gateway.py (claim-before-fire, refuses duplicates, raises UnknownSideEffect for reconciliation), src/continuum/replayguard.py (portable guard), src/continuum/pinning.py and src/continuum/replay_similarity.py (replay correctness), src/continuum/budgets.py (retry caps), src/continuum/recovery/engine.py + src/continuum/recovery/contract.py + src/continuum/recovery/planner.py + src/continuum/recovery/observations.py (max-severity RESUME < ... < ABORT, sealed contract with evidence / reason / next_allowed_action / human_steps), src/continuum/checkpoint/rewind.py (atomic dual-state rewind), src/continuum/analysis/prefix_trust.py (advisory trust). Published checks: docs/recovery_walkthrough.md, benchmarks/fault_injection/ (suite that prints detection_rate / unsafe_resume_rate), src/continuum/benchmark/phase6/ (recovery-correctness suite), docs/RESULTS.md, and the regenerable visual below.

Every row above is traceable to a path that exists on main at the tagged commit. Nothing in this table restates a benchmark number, benchmarks live only in the suite output they already print. See docs/research.md for the full list of published suites and design docs.

Crash recovery, for real

The image below is not a mock. It is the output of python demo-run/generate_crash_visual.py, which runs demo-run/worker.py until os._exit(9) at document 399, calls continuum resume --env dataset=v4 and shows the refusal path (REQUEST_HUMAN, safe:false, exit 21), reconciles the uncertain side effect with a probe, then resumes from the same database and finishes with no duplicate work. The transcript is also saved as docs/assets/crash-recovery.txt for audit.

Regenerate it:

python demo-run/generate_crash_visual.py
# or: python scripts/generate_crash_visual.py

Crash recovery: hard kill mid-batch, refusal, reconcile, resume

Full walkthrough with code is in docs/recovery_walkthrough.md (examples/recovery_walkthrough.py). The minimal bench harness is in references/bench.md (continuum benchmark).

Features

Capability What it gives you
Semantic checkpoints Compact, versioned, inspectable state, not a transcript dump
Idempotent action ledger Refuses duplicate external side effects; surfaces uncertain ones for reconciliation
Environment revalidation Every checkpoint component verified against the current world before resume
Provenance-aware state Agent-reported progress is marked REQUIRES_REVIEW, never self-certifying
Recovery engine Seven recovery modes with a deterministic, sealed next-action contract
Deny-by-default MCP server Thirteen tools, read-only/mutating split, caller allowlist
Framework adapters Generic Python, OpenAI Agents SDK, LangGraph, and LangChain integrations
Secure planning loop Two-signal observation verification escalates high-risk branches to REQUIRES_REVIEW
Periodic revalidation Environment re-checked on a schedule, catching mid-run drift within one cycle
Tamper-evident log Hash-chained event log (52 event types) with integrity verification
Enforcing gate Unclaimed side-effect calls are refused before they fire; deny messages teach the claim protocol
Observation hooks Every file a coding CLI writes becomes digest-verified evidence, outside model control
Session briefing Fresh sessions learn run state deterministically at start, including the last session's reasoning summary
Reconciler probes Registered commands settle uncertain side effects automatically; humans see only the rest
Executable guidance Resume/validate render next steps as runnable commands, not statuses
Enforcing HTTP gateway Outbound calls in any language require claims; responses settle them from reality
OpenTelemetry bridge Tool-call spans from production tracing become evidence with zero code changes
Action index Cross-run idempotency lookups are indexed reads, not full-log scans
Version pinning Caller-asserted prompt/tool/model hashes stored per claim; drift surfaced on resume
Retry budgets Per-action-type attempt caps enforced at claim time; agents see remaining attempts
Multi-agent parent/child Parent resume composes family worst state; uncertain child blocks parent
Informed retry Engine-authored failure summaries injected into post-recovery resumes
Fork semantics Divergent continuations branch into child runs with fresh authority
Log compaction Pre-anchor prefix archived verbatim; live log bounded for month-long runs
Consumed-grant tracking Single-use authority references are marked spent at terminal status; reuse after restore is refused (GRANT_DENIED), defending the checkpoint-restore path against Authority Resurrection
Chain attestation continuum attest signs a run's chain head with Ed25519 so an external verifier can prove history was unaltered as of a known key
HITL dashboard surface Confirm/reconcile/complete buttons with audit parity to the CLI

Security Extension

Two additive security extensions sit on top of the recovery and checkpoint substrate. They do not change resume, replay, or the existing crash-time revalidation path.

Status: specified and tested, not shipped (issue #1030). The modules below are complete and pass their tests, but nothing in the engine or any seam invokes them yet (CONTINUUM revalidates at crash/resume only). Wiring them is a roadmap decision; read the present tense below as the design, not the shipped behaviour.

  • Secure Planning Loop: observations carry provenance and are verified by two independent signals (verified / unverified / contested). A plan branch gated on an unverified or contested observation is escalated to REQUIRES_REVIEW. Decisions are appended to the ledger as PERCEPTION_OBSERVED and BRANCH_RESOLVED events.
  • Periodic Revalidation: reuses the recovery engine on a step interval (default 25) and on app switch, so mid-run environment drift is caught within one cycle instead of only at the next crash.

See docs/PROBLEM.md, docs/RESULTS.md, and STATUS.md.

Empirical Verification

CONTINUUM is verified against real LLM agents, live protocol boundaries, and hard process crashes, not just mock unit tests.

  • Real agents: multi-session Claude Code invoice batches with mid-run SIGKILL, scored 7/7 on mechanics; resumed sessions queried continuum_resume, routed side effects through the two-phase ledger, refused to duplicate verified writes, and respected request_human. Live testing surfaced prompt-drift dedup gaps, closed by canonical path normalization and token-based fallback in ActionLedger.claim().
  • Third-party clients: Gemini CLI and Kilo Code connected over stdio JSON-RPC against the live SQLite store, validating multi-agent co-existence and authorization isolation.
  • Protocol compliance: driven end to end with @modelcontextprotocol/inspector --cli across process deaths; mutating tools deny by default behind CONTINUUM_MCP_MUTATING_CLIENTS; external claims degrade to REQUIRES_REVIEW (safe: false).
  • Self-healing: hard-killed servers recover from orphaned SQLite -wal/-shm sidecars via single-retry cleanup at startup.
  • Scale: roughly 2,864 tests collected (~2,825 passing; ~34 skipped; other outcomes vary by environment) on Python 3.11, 3.12, and 3.13 (unit, hypothesis property-based, concurrency, adversarial). CONTINUUM-Bench runs five crash scenarios plus a dedicated argument-drift scenario, measuring 0 duplicate work and 0 duplicate side effects for CONTINUUM against full duplication for naive replay; a separate 14-scenario recovery-correctness suite (continuum.benchmark.phase6) encodes the crash points from the durable-execution survey as executable assertions, and an 8-fault risk-injection suite (benchmarks/fault_injection/) grades failure handling.
  • Adversarial audit: the full MCP surface was audited over the live protocol; three defects were found and fixed. Method and reproduction steps in test.md.

Horizon-scale benchmark (real runs, no invented numbers)

Generated: 2026-09-30T11:03:24.403017 Horizon scenarios: 5 Passed: 4 Failed: 1

Accuracy: 0.8 Unnecessary escalation: 0.2 Repair precision: 0.8 Duplicate side effects: 0 Duplicate work: 0.0 Compression: 0.138

Scenario Cycles Years Correct Actual Accuracy
horizon_steady_progress_year 141 2.3 resume resume 1.0
horizon_quarterly_drift_year 141 2.3 repair request_human 0.0
horizon_budget_exhaustion_year 141 2.3 request_human request_human 1.0
horizon_compaction_stress_year 177 2.87 resume resume 1.0
horizon_abort_condition_year 141 2.3 abort abort 1.0

Fault-injection: 7 scenarios, detection 1.0, unsafe 0.0

MCP Integration

CONTINUUM ships an MCP server so an agent can record progress, checkpoint, and route external side effects through the ledger without embedding the library:

uv pip install -e ".[mcp]"
CONTINUUM_MCP_MUTATING_CLIENTS=your-client-name continuum-mcp

On Windows, the env-prefix form has no PowerShell equivalent, so set the variable first:

uv pip install -e ".[mcp]"
$env:CONTINUUM_MCP_MUTATING_CLIENTS = "your-client-name"
continuum-mcp

Thirteen tools over stdio. Three are read-only (continuum_validate, continuum_resume, continuum_list_actions); ten mutate. Side effects are two-phase (claim, perform, complete), and mutating tools deny by default behind an allowlist. Agent-reported state is recorded with Origin.EXTERNAL_AGENT provenance and marked REQUIRES_REVIEW.

Verification details, including crash recovery at startup and the end to end Claude Code test, are in references/mcp.md. If a registered server reports CONNECTION_CLOSED, the cause is almost always PATH resolution rather than the server itself: docs/api/mcp.md has the diagnosis and two remedies.

Framework Integration

Nine adapters ship in src/continuum/adapters/ (one in-process facade plus eight integrations), all optional installs so the core stays standard-library-only:

Adapter Class Notes
Generic Python agent GenericAgentAdapter In-process facade; writes trusted (Origin.DETERMINISTIC) state.
Filesystem sandbox FilesystemSandboxAdapter Local directory sandbox, no external service, default for docs and CI.
Python in-process PythonInProcAdapter Runs Python in a temp workdir, records via ledger.
Container ContainerAdapter Docker backed, guarded skip when docker is absent.
Browser BrowserAdapter Playwright backed, guarded skip when not installed.
Kubernetes KubernetesAdapter kubectl backed, guarded skip when not configured.
OpenAI Agents SDK OpenAIAgentAdapter Experimental. Hooks ToolContext / RunHooks; optional openai-agents.
LangGraph LangGraphAgentAdapter Experimental. Wraps a StateGraph; optional langgraph.
LangChain LangChainAgentAdapter Experimental. Drops checkpoint_node into an LCEL Runnable pipeline and the create_agent tool-calling loop; optional langchain.

Each adapter records progress through the ledger and routes external effects through the two-phase intercept/complete protocol. All three framework adapters have end-to-end integration tests and have been driven against a live OpenRouter model, where the runs surfaced and then closed an LLM argument-drift dedup gap and two OpenAI-adapter bugs, including a live hard-crash (os._exit(137) mid-side-effect) proof per adapter. Full usage, live-model results, and runnable examples for every adapter are in references/adapters.md.

Production LangGraph apps can also keep their native persistence API: make_continuum_checkpointer(storage) implements LangGraph's BaseCheckpointSaver over CONTINUUM's storage, so every put lands in the same hash-chained, provenance-tagged event log (see references/adapters.md).

Three further production frameworks are covered by thin, SDK-free hook surfaces in adapters/thin.py:

Framework Interception surface Entry point
CrewAI global before/after tool-call hooks install_crewai_hooks(storage, run_id)
AutoGen core FunctionTool.run_json wrapped in place wrap_autogen_tool(tool, storage, run_id)
Pydantic AI async Hooks capability Agent(capabilities=[wrap_pydantic_ai_hooks(storage, run_id)])

For stacks none of these reach: continuum gateway enforces claims on outbound HTTP from any language, continuum.otel.make_span_processor(storage) turns existing OpenTelemetry tool spans into evidence, and continuum serve exposes the same operations as the MCP tools over a language-agnostic JSON wire protocol (stdio, or HTTP via --transport http with CONTINUUM_SERVE_TOKEN auth).

Resuming agent- or MCP-reported runs

State reported over MCP, or through the OpenAI adapter, carries Origin.EXTERNAL_AGENT provenance and resolves to request_human until confirmed. LangGraph and LangChain runs use Origin.DETERMINISTIC and resume directly. To clear review and resume:

continuum confirm <run_id>   # records REVIEW_CONFIRMED, then re-assesses
continuum resume <run_id>    # now reports RESUME

Over MCP the equivalent is the continuum_confirm tool followed by continuum_resume. Confirmation is a one-time, human-attested event: the escape hatch for the self-certification safety, so an externally-driven run is never permanently stuck.

Core Concepts

The deep reference for each concept lives in references/concepts.md.

  • Semantic Checkpoints - a compact, versioned representation of what the agent needs to continue.
  • State Validation - every component independently verified; staleness propagates through the dependency graph.
  • Idempotent Action Ledger - external side effects tracked and de-duplicated; uncertain outcomes raise instead of silently retrying.
  • Recovery Modes - RESUME, REPAIR_AND_RESUME, ROLLBACK, WAIT, REQUEST_HUMAN, ABORT (plus REPLAN).
  • Recovery Contract - a deterministic, integrity-sealed, gated next action.

Architecture

CONTINUUM is organised around one invariant: every fact carries its origin, and trust is earned, never assumed. Why this matters for a startup: an agent that runs for weeks must not lose work when its context is lost, and it must not waste tokens, cost, or fire a tool twice.

System at a glance: universal adapter, one log, any harness

Any harness plugs into the same hash chained log. The same run can be written by Claude Code, resumed by LangGraph, inspected by the CLI, and approved on the dashboard. No framework cooperation is required.

  Claude Code ─┐
  Gemini CLI ──┤
  Codex ───────┤
  LangGraph ───┼── 5 seams ──►  One durable log  ──►  Recovery + Dashboard + CLI
  LangChain ───┤                (hash chained,        (sealed contract,
  OpenAI SDK ──┤                 provenance tagged,     verify, health,
  CrewAI ──────┤                 exactly once)          family)
  Any HTTP ────┤
  Any OTel app ┘

  Seams: 1 In-process  2 MCP  3 CLI hooks  4 Gateway  5 OTel

The three guarantees (the demo proves each one)

  1. No self-certification. Agent reported state is EXTERNAL_AGENT and degrades to REQUIRES_REVIEW until a human REVIEW_CONFIRMED. Only trusted writers produce DETERMINISTIC state.
  2. Side effects require claims. Every external effect is claimed in an idempotent ledger before it fires. Unclaimed effects are blocked at the boundary, duplicates are refused, uncertain outcomes raise for reconciliation.
  3. Recovery verifies against reality. Resume checks file digests, dependency versions, and model identity before saying safe. Staleness propagates dependency -> evidence -> finding -> decision plus PlanStep.depends_on so only affected steps repair.

Five integration seams

Seam How to connect What it gives you
1 In-process GenericAgentAdapter.intercept_action(...) and wrap_tool(key_fn=...) on LangChain, LangGraph, OpenAI Agents SDK Python frameworks, trusted writes
2 MCP server continuum-mcp 13 tools over stdio (continuum_record_progress, continuum_intercept_action, continuum_complete_action, etc.) Any MCP capable client, 3 read-only + 10 mutating, allowlist CONTINUUM_MCP_MUTATING_CLIENTS
3 CLI lifecycle hooks continuum hooks install claude-code --with-gate also gemini and codex Coding CLIs: SessionStart briefing, PostToolUse observe, PreToolUse gate; no CLAUDE.md needed
4 Enforcing HTTP gateway continuum gateway --port 8765 with .continuum/gateway.json Any language, any outbound HTTP must have a claim, gateway settles from real status code
5 OpenTelemetry bridge make_span_processor(storage) Any traced app, spans become TOOL_COMPLETED evidence

Thin hook surfaces for CrewAI, AutoGen, Pydantic AI live in adapters/thin.py with no SDK required.

Enforcement pipeline: why no duplicate and no invalid call

The gate to observe pipeline closes the gap at the harness boundary. This is what saves tokens and cost and blocks invalid tool calls.

PreToolUse hook                    PostToolUse hook
    |                                    |
    v                                    v
continuum gate                    continuum observe
    |                                    |
    |-- no claim? DENY (exit 2)          |-- TOOL_COMPLETED event:
    |   + instructions to claim          |     path, bytes, sha256 on disk now
    |                                    |
    |-- live claim? ALLOW                |-- disk checked status:
    |                                    |     verified / changed / missing
    v
agent performs effect
    |
    v
continuum_complete_action  (settled from reality, not from report)
    |
    v
ledger marked COMPLETED: next replay returns cached result, not a second fire

Unknown host is denied fail closed, not an open relay. Shell Bash/curl is the documented v1 blind spot.

Recovery decision tree: weeks until done, correct and exactly

The engine takes the most cautious signal, so safety never loses to convenience.

RESUME < REPAIR_AND_RESUME < REPLAN < WAIT < REQUEST_HUMAN < ROLLBACK < ABORT

Every continuum resume returns a sealed contract with: recovery status and safe, verified/invalidated components, executable human_steps (exact shell to run), post checkpoint observations disk checked, pinning drift, and family aggregation for multi agent continuum tree. Briefing continuum briefing injects that contract on every fresh claude SessionStart, so saying hi after you killed the terminal resumes from the last good prefix.

Why this saves tokens, cost, and invalid calls

  • Tokens: Semantic checkpoint stores Goal + Plan + Progress not a transcript dump. Briefing serves only verified state plus a 4096 cap reasoning summary, not the error tail that self conditioning shows degrades the next session. Informed retry recovery/summary.py injects an engine authored summary, not raw history.
  • Cost: Ledger action_index refuses duplicate side effects even under argument drift like relative vs absolute paths (invoice:INV-001 stable key), so the same API is not paid twice after a resume. Budgets budgets.py cap retry storms at claim time. continuum benchmark prints 0 duplicates for continuum vs 50 for naive.
  • Invalid calls: Gate and gateway and replayguard langgraph_protected_node block unclaimed or replayed tool calls before they fire. Pinning pinning.py surfaces prompt or tool drift on resume.

Storage architecture

Schema v6. SQLite is primary, Postgres is CI verified. One log, many projections.

Table Purpose
events Hash chained append only log (52 event types)
runs Run metadata with parent_run_id for multi agent
versions SemanticState snapshots per checkpoint
checkpoints Sealed checkpoint records with RECOVERY anchors
action_index Cross run idempotency projection (schema v3+): indexed reads, not full scans
events_archive Compacted prefix storage (schema v5+): continuum compact bounds live log for weeks
lg_checkpoints / lg_writes LangGraph native persistence (schema v4+): make_continuum_checkpointer(storage)

Module map: one library, many surfaces

CONTINUUM is one library (src/continuum, 132 modules) plus a large test suite (187 test files, ~2,906 tests). All modules append to and replay one hash chained event log:

Module Role
events.py Append only, hash chained event log and verify() trusted_through
state/ Projection project(), validation, extraction: staleness propagation
storage/ SQLiteStorage v6, postgres.py, migrations.py, actionindex.py
actions/ Idempotent ledger claim/complete/reconcile, idempotency.py key + canonicalization + token fallback, consumed grant tracking GRANT_DENIED
checkpoint/ Policy driven checkpoints manager.py policy.py with RECOVERY anchors and prune
recovery/ Engine, planner, sealed contract contract.py, guidance human_steps, observations disk checked, family rollup, fork semantics, summary informed retry
gate.py Pre tool use enforcement: allow or deny against ledger claims
gateway.py Enforcing HTTP proxy: claim before fire for outbound requests
replayguard.py Portable guard: evaluate, protected_call, langgraph_protected_node: closes ACRFence replay hazard
hooks.py clienthooks.py Shared checkpoint hooks and installer profiles claude-code gemini codex
budgets.py Retry budget registry and evaluation per action type
pinning.py Version pinning normalisation and drift detection on resume
replay_similarity.py Semantic similarity backends exact/fuzzy/embedding for replay vs fork
reconcilers.py Probe registry .continuum/reconcilers.json for automatic settlement
adapters/ 9 class adapters + thin hooks thin.py CrewAI AutoGen Pydantic AI + LangGraph store
mcp/ 13 stdio tools plus authz authz.py token auth, allowlist, confirmation token
serve/ Sidecar stdio JSON wire + HTTP CONTINUUM_SERVE_TOKEN
dashboard/ Web dashboard app.py hitl.py with HITL buttons confirm/reconcile/complete, prefix trust advisory, pins
cli/ 49 argparse commands, exit codes as verdict: runs, start, inspect, resume, verify, health, tree, benchmark, attest, dashboard
otel.py OpenTelemetry span processor bridge
benchmark/ CONTINUUM-Bench harness: 5 crash scenarios + argument drift + 14 scenario recovery suite + 7 fault risk injection

Honest limitations

  • Gate does not see inside shell commands (Bash/curl bypass structured tool claims)
  • Postgres backend is CI tested but not battle tested in production
  • One level of multi agent hierarchy v1
  • Large payload offloading (#254) not yet implemented
  • Weeks scale benchmark with token cost table lands in #550 board (#568 to #570)

Full reference in references/architecture.md. And the months plane that builds on this (provenance causal graph, authority resurrection, admissibility, liveness) is pinned as board #550 with 20 sub issues #551 to #570.

API and CLI

Python surface (EventType, Run, SQLiteStorage, diff_states, project) and the adapter API are documented with runnable examples in references/api.md. The CLI is the same surface in shell form:

continuum runs                                   # list runs
continuum inspect <run_id>                       # semantic state
continuum validate <run_id> --env dataset=v4     # validate, read-only
continuum resume <run_id> --env dataset=v4       # recovery decision + contract + next steps
continuum checkpoint <run_id>                    # force a checkpoint, mutates
continuum actions <run_id>                       # external side effects
continuum reconcile <run_id>                     # settle uncertain effects with probes
continuum complete <run_id>                      # close a run as done, from the keyboard
continuum verify <run_id>                        # re-audit the event hash chain
continuum budget <run_id>                        # retry-budget usage per action type
continuum compact <run_id>                       # archive pre-anchor log prefix
continuum tree <parent_run_id>                   # show parent + children with recovery states
continuum attest <run_id> --key signer.pem       # sign the chain head for an external verifier

All wiring is host-side; the model's cooperation is optional:

continuum hooks install claude-code --with-gate   # coding CLIs: evidence, briefing, gate
continuum gateway --port 8765                     # enforcing HTTP proxy for everything else
provider.add_span_processor(continuum.otel.make_span_processor(storage))  # OTel to evidence
continuum-mcp                                     # anything MCP-capable: the twelve-tool server
continuum briefing                                # session-start context injection
continuum budget <run_id>                         # retry-budget usage report
continuum tree <parent_run_id>                    # multi-agent hierarchy view

Optional registries live beside your code and are data, not code: .continuum/gate.json (side-effect tools + stable-key templates), .continuum/reconcilers.json (probes that check external systems), .continuum/gateway.json (upstream routes).

Most commands accept global --json before the subcommand (e.g. continuum --json resume RUN), and read-only commands never write, so they are safe against a live database while an agent is mid-run. The two exceptions are audit-only and fire solely for a run that is already parked: resume records its webhook delivery outcome when the verdict is REQUEST_HUMAN and an endpoint is subscribed, and watch records a liveness verdict flip; neither append projects onto state or moves the exit code. Exit codes are a safety contract (only a verified-safe run exits 0). Full command list, exit-code table, and state-diff output in references/cli.md.

Roadmap

Phase Component Status
1-11 Data models, semantic state, persistence, checkpointing, validation, action ledger, recovery engine, CLI, crash-recovery examples, environment snapshots/diffs, framework adapters Complete
12 Benchmark suite (CONTINUUM-Bench) Complete (minimal harness)
13 Cloud API (FastAPI + PostgreSQL) Partial: the PostgreSQL storage backend and the HTTP sidecar transport (continuum serve --transport http) are shipped and CI-tested; the hosted multi-tenant service is not started
14 Dashboard Complete (continuum dashboard)
15+ Enforced durability: observation hooks, gate, session briefing, reconciler probes, enforcing gateway, OTel bridge, action index, executable guidance, multi-client installers, semantic replay detection, version pinning, retry budgets, log compaction, HITL surface, fork semantics, informed retry, multi-agent aggregation Complete (see issue #213)
Next Months-scale durability plane: milestone-anchored plans (#312), structured attempt memory (#313), atomic dual-state rewind (#292), public recovery-correctness benchmark (#293) Planned (draft spec in docs/UPGRADE_SPEC.md); webhook-out notifications (#305) shipped

Beyond the original plan: the MCP server, MCP authorization and caller-authentication layers, provenance and anti-self-certification, community files, schema versioning with forward migrations, a bounded recovery context, consumed-grant tracking, Ed25519 event-chain attestation, the native LangGraph checkpointer, and wheel artifacts on every push to main are shipped. See STATUS.md for the verified-vs-believed breakdown and open correctness bugs.

What CONTINUUM Is Not

Not this This instead
An LLM A reliability layer for agents that use LLMs
An agent framework A recovery layer that plugs into any framework
A vector database Structured semantic state, not embeddings
A RAG system Verified checkpoints, not retrieval-augmented memory
A workflow engine A recovery layer, not an orchestrator

The core abstraction: semantic state + environment validation + action reconciliation = safe recovery.

The verdict is advisory, not enforced

CONTINUUM tells you the truth about a run. It does not stop your process when the answer is unwelcome. RecoveryDecision.permits() reports what the contract allows; nothing in the library intervenes if a caller ignores a False and acts anyway. A worker that calls CheckpointManager.restore directly can continue past a REQUEST_HUMAN verdict, because it never asked the engine whether to.

This is a deliberate contract for a library: forcing enforcement inside a call the caller made to inspect a verdict would surprise the callers who legitimately want to read it and then override it.

Enforcement exists, but as separate seams you opt into. None is enabled by a plain pip install continuum-agent:

Seam How to enable
Host gate continuum gate
HTTP gateway continuum gateway
Replay guard for framework calls continuum.replayguard in-process
Observation hooks continuum hooks install

The one enforcement that ships enabled is the CLI exit code. continuum resume exits non-zero unless the run is verified safe (RESUME), so continuum resume "$RUN" && ./start-agent.sh cannot launch onto stale state. Every other mode maps to a distinct non-zero code, and a mode nobody has classified falls through to UNSAFE rather than OK.

If you want the verdict enforced, wire a seam or gate your pipeline on the exit code. Do not assume the library is supervising the process.

Related work

CONTINUUM sits at the overlap of durable execution, idempotent side-effect tracking, and crash recovery for LLM agents. The closest neighbors are machine-checked resume contracts (Khan 2026), agentic transaction processing with constraint-gated admission (Mnemosyne 2026), checkpoint-rollback attack analysis (ACRFence 2026), and design-level prompt-injection defense (CaMeL 2025). The full annotated list, foundations, and citation audit are in references/related-work.md.

Status and limitations

  • Tested: 1,360 passed + 23 skipped in a full run at the 2026-08-24 audit of this tree; CI enforces the suite on Python 3.11, 3.12, and 3.13, and counts vary by platform and optional services such as Postgres (see STATUS.md). The MCP surface has also been audited adversarially over the live protocol; see test.md.
  • On PyPI as continuum-agent 0.1.2 (pip install continuum-agent; clone still works via pip install . see Quick Start).
  • MCP caller authentication is opt-in per deployment. When CONTINUUM_MCP_TOKEN is set, the server refuses every mutating tool unless the caller presents that shared secret in the initialize handshake's _meta.authToken; per-caller secrets are available via CONTINUUM_MCP_CLIENT_TOKENS (name:secret pairs). Without any token configured, authorization is by declared identity only (the historical default, preserved for local single-user use).
  • Confirming self-reported state over MCP requires a separate secret. continuum_confirm refuses every caller until the operator sets CONTINUUM_MCP_CONFIRM_TOKEN, because an agent allowed to record progress must not also be able to confirm it. The default path stays human-driven: run continuum confirm <run_id> on the host.
  • Unbuilt components: Cloud API (Phase 13).
  • Shell command enforcement gap: the gate enforces claims for structured tool calls but cannot see inside Bash/curl commands. Documented as v1 scope refusal.
  • Framework adapters remain experimental. All three framework adapters now carry live-model soft-resume and hard-crash proofs (OpenRouter, gpt-4o-mini), including the crash contract that blocks resume on an uncertain side effect, and now have crash-and-resume verification tests achieving parity with the generic facade (Refs #285). Prefer GenericAgentAdapter for production recovery.
  • Agent/MCP runs need an explicit confirm before auto-resume. Externally-reported state is REQUIRES_REVIEW, so continuum resume returns request_human until a human confirms. By design, not a bug; see Framework Integration.
  • e2e autonomy test series (issue #6): three full Claude Code runs scored 7/7 mechanics with unprompted recovery behavior observed. Further iterations across diverse prompt styles remain open.

About

In early 2026 I saw long running agents fail on recovery, not reasoning. Checkpoints were treated as proof to continue, not evidence to verify. Surveying Temporal, LangGraph, ACRFence 2603.20625 and self conditioning 2509.09677, I found the gap was a portable verification substrate that asks, given the state at time T and the world as it is now, is it still safe to continue.

Over three weeks I built CONTINUUM from one invariant, every fact carries its origin. The result is a hash chained log with verify(), a ledger with stable key deduplication, a gate and gateway that block unclaimed effects, and a recovery engine that seals a contract. Five seams expose the same log to Claude Code, LangGraph, LangChain, OpenAI, HTTP and OpenTelemetry. Validated with real kills and ~2,864 tests, it prints 0 duplicates where naive replay prints 50.

CONTINUUM was created by Anandhu P Shaji (@Cyrax321 · LinkedIn) and is maintained by the original creator. It is open source under the Apache-2.0 license. Community contributions are welcome via CONTRIBUTING.md and are credited in AUTHORS.md and graphs/contributors.

Contributing

This project is open source under Apache 2.0 and deliberately built to be extended: by researchers validating the recovery semantics, by engineers porting the ledger or MCP server to other frameworks or languages, and by anyone turning the planned roadmap into reality. A good place to start is the good first issue label on the issue tracker, or the open correctness bugs listed in STATUS.md.

Open an issue before submitting large PRs. See CONTRIBUTING.md for the full contribution guide, including the Code of Conduct.

Contributors

Thanks to our community contributors, ordered by contributions: @Adhi1-2, @abyyxhek, @Amiirhosseini, @yuki-fuyutsuki, @vjymisal0, @adity982, @dchaudhari7177, @tasodoufu, @anya-research, @lesbass, @Parthipashok04, @Samearth17, @stoppo22, @timothyanderson096-ocdealcheck, @unmoha, @aastha-m22, @as950118, @asarakhatun17-lgtm, @challenge456, @gouthamkrishnak2003, @ItzSaurav, @mhaye9545, @Newer1107, @okestroHjJeong, @quangshuynh, @Rahul-pamula, @Shaisolaris, @VedantMadane, @zynx-real.

Sponsor

If CONTINUUM helps your agents recover reliably, consider sponsoring to support long term maintenance.

Sponsor Cyrax321

Become a sponsor: GitHub Sponsors, or add FUNDING.yml custom link if you prefer another platform.

License

Apache 2.0 - see LICENSE.


Deep reference material:

Quick links

Search skills and MCP servers

Fuzzy search across 23,137 skills and servers