Agent Skills

podcli

Open-source AI podcast clipper. Generate vertical clips with face tracking and burned-in captions. CLI, MCP server, web app.

Install

npx -y podcli
README.md

podcli

Open-source AI podcast clipper.
Turn a long episode into short clips with face tracking and burned-in captions. Drive it from the CLI, a web studio, or your coding agent.

podcli.com · Docs · Install · MCP

license: AGPL-3.0 stars

Podcli demo

▶ Watch with sound on X

podcli process episode.mp4

That one command transcribes the episode, picks the moments worth clipping, crops to whoever is speaking, and burns the captions in. Transcription and rendering run on your machine. The only network calls are the optional Claude or Codex requests when you use AI clip scoring.

Install

No prerequisites. The installer fetches a self-contained binary, and the first run provisions Python, Node, FFmpeg, whisper.cpp, and the models it needs into a managed folder.

macOS and Linux

curl -fsSL https://podcli.com/install.sh | sh

Windows (PowerShell)

irm https://podcli.com/install.ps1 | iex

Runs on macOS (Apple Silicon), Linux (x64 and arm64), and Windows (x64). Intel Mac support is in progress.

Quick start

podcli                       # interactive menu, opens the web studio
podcli process episode.mp4   # transcribe, pick moments, render clips
podcli multicam ~/ep12       # sync every camera and mic, cut to the speaker, render the episode

Clips land in podcli-clips/ in the directory you ran it from, so each show keeps its own renders. Everything else (knowledge, presets, assets, clip history, cache) lives in one managed folder that follows you between directories. Set PODCLI_OUTPUT to render somewhere fixed instead.

What you get

Clips

  • 9:16, 16:9, or 1:1, with captions sized for each canvas
  • Face tracking that follows the speaker, split-screen layouts included
  • Multi-segment cuts that drop filler, long pauses, and tangents
  • Four caption styles: branded, hormozi, karaoke, subtle
  • Logos, intros, outros, and background music from a reusable asset library
  • Loudness-normalized audio and hardware encoding on VideoToolbox, NVENC, and VAAPI, with a CPU fallback

Editing the full episode

  • Multicam edit: point at one folder of camera and mic files, confirm who each file belongs to, and podcli syncs them by audio and corrects clock drift between devices
  • Automatic camera cuts that follow whoever is talking, with wide shots for crosstalk and a guest's answer held on the guest
  • Remote shows too: one file per person becomes a split screen, and a single gallery recording is split into a camera per tile. Questions stay on the split and the guest goes full frame on long answers
  • Render the episode as an MP4 with one WAV per person. Studio shoots can also go to Premiere or Final Cut as a timeline that points at your original files
  • Same edit from the studio (Multicam edit), the CLI (podcli multicam <folder>), or an agent (manage_multicam). Re-running reuses the sync and the cut, so only the step you changed runs again
  • With podcli Pro, steer the cut in the cloud editor: podcli multicam <folder> --cloud sends small previews and the transcript (your camera files stay on your computer), and --pull renders the cut you made there

Finding the moments

  • Whisper transcription with speaker diarization, or bring your own transcript as .txt, .srt, or .vtt
  • AssemblyAI as an alternative engine, and yt-dlp to pull an episode straight from a URL
  • AI scoring against your knowledge base, checked against your episode database so it stops resuggesting moments you already published
  • Audio energy and laughter detection to build highlight reels

The studio at localhost:3847

  • Library, episode workspace, multicam edit, per-clip detail, highlights, thumbnails, content, analytics, assets, knowledge, config, integrations, and MCP setup
  • ⌘K command palette across pages, clips, and assets
  • Titles, descriptions, tags, and hashtags, with any section regenerated on your own guidance
  • Thumbnail studio for 16:9 and 9:16, with frame and text options
  • Transcript corrections that carry through to every render

Shipping it

  • 27 MCP tools, so an agent can transcribe, score, render, and publish through conversation
  • YouTube publishing plus performance analytics to see which clips landed
  • DaVinci Resolve export as FCPXML when you want to finish by hand
  • Presets, clip history with duplicate detection, and a transcript cache

Why podcli

If you are weighing podcli against the cloud clippers, this is the difference:

  • Runs locally. Transcription and rendering happen on your machine by default, so episodes stay there. Only the optional cloud engine (AssemblyAI) and publishing to YouTube send anything out.
  • Free and open source under AGPL-3.0. Exports are unlimited, full quality, and watermark-free.
  • Agent-native. 27 MCP tools let Claude Code or Codex drive the whole flow, transcription through publishing.
  • A knowledge base keeps titles, captions, and descriptions in your show's voice, and stops the engine from resuggesting moments you already published.
  • DaVinci Resolve handoff. Export any clip as FCPXML when you want to finish the edit yourself.

Use it from your agent

podcli is an MCP server, so an agent can transcribe, suggest clips, and render them through conversation.

podcli mcp install    # registers it with Claude Code

Claude Desktop and Codex setup is in the MCP docs.

Content workflow

PodStack ships with podcli as a set of Claude Code slash commands. They take a transcript to a publish-ready package: scored moments, titles, descriptions, thumbnail briefs, a brand review, and a publish checklist.

/produce-shorts

The commands live in .claude/commands/. CLAUDE.md describes each one.

Docs

Guide What's in it
Getting started Install, first episode, the whole flow
The studio Web UI: library, episodes, content, highlights
CLI Commands, flags, presets, assets
MCP server Agent setup and available tools
Captions and formats Styles, aspect ratios, cropping
Configuration Environment variables, config profiles, transcript format

Docs are open source at nmbrthirteen/podcli-docs.

Contributing

See CONTRIBUTING.md for the dev setup and conventions, and RELEASE.md for how releases are cut.

Credits

Content workflow powered by PodStack, inspired by gstack by Garry Tan.

License

AGPL-3.0. See LICENSE.

Need podcli without AGPL terms? A commercial license is available. Email siradze@nikusha.me with a one-line description of your use case.

Search skills and MCP servers

Fuzzy search across 23,137 skills and servers