Agent Skills

Orchestrate a video run rendered by VEED's engine — stylized captions over footage, edits and reframes, layered motion graphics, or graphics with no footage at all. Takes any number of source videos, including none. Use when the user wants video made, edited, or captioned by an agent.

Install

npx skills add https://github.com/veedstudio/open-edit --skill open-edit
SKILL.md

open-edit

You author the piece yourself, as an HTML page, and render it with render. OpenEdit supplies what a bare shell does not: VEED's hosted services on one login, measured tools for cutting and sound, and a renderer that re-renders only the stretch you changed. openedit below is short for npx @veedstudio/openedit-cli; every command explains its flags with --help.

Setup

Once per workspace, before the first command. WORKSPACE is the current project root, or the current directory outside a project. It is done when your own init run below ends on ready — OPEN_EDIT_ROOT=…, or when a session-opening note says Open Edit preflight is ready and OPEN_EDIT_ROOT=<path>. No further preflight is needed; a note that lists APPROVAL REQUIRED never means that, whatever else it quotes.

npx --yes @veedstudio/openedit-cli init --dry --workspace "$WORKSPACE"
npx --yes @veedstudio/openedit-cli init --workspace "$WORKSPACE"

Read the final preflight: line, not the exit code: ready — OPEN_EDIT_ROOT=… means go, and runs/ and the recorded preferences live under that root. If init prints APPROVAL REQUIRED, tell the user every exact action and wait for an explicit yes, then run it again with --auto-approve. When the final line says the user runs the install commands, --auto-approve cannot run them: give the user those commands and run init again once they are done. Never install anything machine-global without that yes, and never infer it from the render request.

One folder per piece

Keep everything a piece is made of under runs/<key>/: the page, its assets, renders, audio. The code is the edit: that folder is how the user comes back later to change one scene or reuse the piece for another video, so nothing it needs lives anywhere else.

Tools

Need Command
A transcript with real per-word times transcribe <video> [...] runs the provider recorded in $OPEN_EDIT_ROOT/.open-edit-prefs.json. None recorded (the command refuses and says so), "No VEED login found", or a failed run: read TRANSCRIPTION.md first
Remove, reorder or join takes read CUT.md (speech-probe, apply-edl, retime-transcript)
Join whole clips whose shapes disagree concat-videos (re-encodes to one canvas at 30 fps)
Look at a video, a page or a pile of images frames <video> --at <sec,...> --sheet (one sheet of those moments; --images for a folder of references), render --stills
Cut the subject out of its background background-removal (the free VEED route; when it cannot run, the command stops with fal's price, and only --fal buys)
A presenter speaking a script Fabric: read FABRIC.md first
New audio on a face lipsync
Any fal model: images, video, music, voice, effects fal schema <model>, then fal run <model> --input <json or @file> --run runs/<key> (--batch <jobs.json> for many); local files in the input are uploaded for you
Web fonts as local files fonts <Family> [...] --out <dir>
Real, licensed pictures stills search, stills show, stills save
One soundtrack from several pieces mix-audio runs/<key> (voice tracks duck the bed; mix-audio --help has the spec)
Sound on a render, at delivery loudness mux-audio --video <render> --audio <source clip or built track> --out <file>
The piece in VEED's editor for the user to change, or a project they bring from it read VEED.md (veed-project, veed-pull)

Fabric, background removal, lipsync and veed-project reuse the VEED transcription login: same veed.io account, same stored token, no second sign-in. The editor hand-off's bookmarks run in the user's own VEED session in the browser. background-removal --fal needs no VEED login at all.

Render

openedit render runs/<key>/index.html --out runs/<key>/out.mp4 --fps <fps> --duration <seconds>

The canvas is the page's #stage (or [data-stage]) element, captured at its own size, else 1920x1080. --width/--height instead capture that much of the page from its top-left corner and ignore the stage. With footage, the canvas and frame rate are the source's own, and frames prints both. Pass its frameRate (such as 30000/1001) to --fps, never the decimal beside it.

  • The renderer owns time. It drives a virtual clock (performance.now, Date.now, requestAnimationFrame, timers, a seeded Math.random), sets every CSS and Web Animations animation to the frame's time and seeks GSAP's global timeline. Anything else that moves with time goes in window.__seek(t) (it may be async).
  • It waits for fonts, images and videos before the first frame, and shows each <video> at the exact frame. Keep fonts and assets as files beside the page.
  • --from <s> --to <s> re-renders only that stretch and reuses the rest. After a change, re-render what changed, not the piece.
  • --stills 0.5,2,4 --sheet <file> shows moments without a full render. --transparent gives alpha.
  • It fails loudly on a page error, a blank render or a Chrome that did not start; read the reason it prints. A load it lists as failed (a video, an image, a font) is a wrong render: fix it and re-render before delivering.
  • The output is silent: mux-audio lays the sound on (build a many-piece track first with mix-audio).

No footage

If the user attached no video, read the ask before assuming one is needed:

  • A clip is coming: wait for it.
  • A talking head from a script: Fabric, per FABRIC.md.
  • A clip from another model (anything on fal, their key and their bill): say in the same breath that captions are transcribed from the clip's speech, and most generators return silent clips.
  • No video at all: build from stills, graphics, generated imagery and audio. That is a whole run.

If it is unclear which, ask.

Money

  • VEED transcription consumes the VEED transcription credits of one workspace. On an account with several, the CLI will not pick: name the one the user chose with --workspace.
  • Fabric consumes the AI Playground credits of the workspace chosen for it; FABRIC.md has its three stops.
  • Every fal call (fal run, lipsync, background-removal --fal or --fast) bills the user's own fal account. Before the first paid call, say in one line what you will make and what it costs, and let them answer. Background removal's default route is free.
  • Report what was spent, and on which account, when you deliver.

Talking to the user

One line when you start, then the deliverable's path and a sentence or two on the result. No step-by-step progress and no internals: run keys, command names, raw tool output. Pass on in plain terms a warning about the result, such as words transcribed without timings. If you must stop, say in plain terms what is wrong on screen and what the options are.

Related skills

video-editgenmedia-labs715KEdit existing video on RunComfy — this skill is a smart router that matches the user's intent to the right edit model in the RunComfy catalog. Picks Wan 2.7 Edit-Video (general restyle / background swap / packaging swap, identity + motion preservation), Kling 2.6 Pro Motion Control (transfer precise motion from a reference video to a target character), or Lucy Edit Restyle (lightweight identity-stable restyle / outfit swap). Bundles each model's documented prompting patterns so the skill gets shai-video-generationgenmedia-labs714KGenerate AI videos on RunComfy via the `runcomfy` CLI — a smart router across the full video-model catalog: HappyHorse 1.0 (Arena #1, native in-pass audio), Wan-AI Wan 2-7 (open weights, audio-driven lip-sync), ByteDance Seedance v2 / 1-5 / 1-0 (multi-modal cinematic), Kling 3.0 / 2-6, Google Veo 3-1, MiniMax Hailuo 2-3, ByteDance Dreamina 3-0. Covers text-to-video (t2v), image-to-video (i2v), and Veo's video-extend endpoint. The skill picks the right model for the user's intent (Arena-#1 qualitai-musicgenmedia-labs714KGenerate AI music on RunComfy via the `runcomfy` CLI — a smart router across the music-model catalog. Routes to ElevenLabs AI Music Generation (premium 44.1 kHz stereo vocal tracks, 5 s–5 min, $0.0083/s) and ACE Step / ACE Step 1.5 (StepFun-AI open-weights, tag-driven composition, multilingual lyrics, $0.0002–0.0003/s, ~27× cheaper), plus ACE Step audio-inpaint (regenerate a time range inside an existing track) and ACE Step audio-outpaint (extend a track before or after). Picks the right model fimage-to-videogenmedia-labs713KAnimate any still image on RunComfy — this skill is a smart router that matches the user's intent to the right i2v model in the RunComfy catalog. Picks HappyHorse 1.0 I2V (Arena #1, native audio, identity preservation) for general animations, Wan 2.7 with `audio_url` for custom-voiceover lip-sync, or Seedance 2.0 Pro for multi-modal animation from image + reference video + reference audio. Bundles each model's documented prompting patterns so the caller gets sharper output without burning iterat

Search skills and MCP servers

Fuzzy search across 23,137 skills and servers