Agent Skills

video-transcription

Transcribe speech in local or remote videos with timestamps. Convert existing timed transcripts locally into SRT or ASS; use media-analysis for visual understanding.

Install

npx skills add https://github.com/postplusai/postplus-skills --skill video-transcription
SKILL.md

Video Transcription

Use When

  • The input is a video file and the goal is speech extraction, timed transcript, caption generation, multilingual transcript, or edit-prep timestamps.
  • Use media-analysis instead when the user needs semantic visual analysis.

If a timed transcript already exists and only subtitle output is requested, skip transcription and read the local subtitle conversion reference.

Do Not Use When

  • The task needs new creative generation or visual analysis rather than speech or subtitles.
  • Required inputs are missing and guessing would change the result.

Execution Boundary

  • Hosted video transcription runs through the public postplus media transcribe verb and is async. The generated example below shows the endpoint key.
  • Pass a local path, HTTPS URL, existing PostPlus media reference, or data URI directly to --video. The CLI validates and prepares local media before the single hosted submit.
  • Request timestamps by default when results drive subtitles or edit decisions.
  • Hosted video transcription is async. Submit records the run handle, current status, and transcript artifacts when available. Inspect the returned format; do not assume a fixed normalized transcript schema.

Source And Path

  • Before submit, derive durationSeconds from the source video or URL and pass it through the endpoint's duration flag for request validation.
  • Start with one source file before larger batches.
  • Keep internal requests, responses, normalized transcripts, and downloaded artifacts under .postplus/video-transcription; keep final user-facing transcript exports outside .postplus.

Handoff

  • If status is pending, preserve the result path and follow the CLI-returned action or resume command for the same operation. Do not submit another job. Stop and report when the CLI wait/recovery boundary is reached.
  • When SRT/ASS is requested, use the actual timed transcript and read local subtitle conversion. Convert locally without another hosted request; do not invent a CLI export command.

Stop Conditions

  • Stop when required user intent, source evidence, or owned input artifacts are missing and guessing would change the result.

Public Command Boundary

  • Choose the smallest matching command or workflow from the user input and run it directly.

  • Readiness diagnostics: postplus doctor --skill video-transcription.

  • Use postplus media schema --json only when you need the full endpoint, flag, and enum contract or are repairing an unknown request shape.

  • Run the hosted transcription job with the generated command below; do not use another execution interface.

  • Pass the source directly through --video; do not pre-upload it or construct a manual request object.

  • If the CLI returns a quote-confirmation challenge, obtain user approval for its scope and cost before running postplus quote confirm --json --challenge-file <challenge.json> and retry with the returned token.

postplus media transcribe transcription-video \
  --video ./reference.mp4 \
  --duration-seconds 1 \
  --wait \
  --output ./result.json

Follow the CLI's structured result and reported next action; do not infer recovery from free-text messages. Wait for explicit user approval when requested; an action does not authorize spending, publishing, or overwriting. Resume the same operation through its returned checkpoint or action; never resubmit uncertain work, repeat exhausted recovery, or switch providers to bypass failure.

Related skills

video-editgenmedia-labs715KEdit existing video on RunComfy — this skill is a smart router that matches the user's intent to the right edit model in the RunComfy catalog. Picks Wan 2.7 Edit-Video (general restyle / background swap / packaging swap, identity + motion preservation), Kling 2.6 Pro Motion Control (transfer precise motion from a reference video to a target character), or Lucy Edit Restyle (lightweight identity-stable restyle / outfit swap). Bundles each model's documented prompting patterns so the skill gets shai-video-generationgenmedia-labs714KGenerate AI videos on RunComfy via the `runcomfy` CLI — a smart router across the full video-model catalog: HappyHorse 1.0 (Arena #1, native in-pass audio), Wan-AI Wan 2-7 (open weights, audio-driven lip-sync), ByteDance Seedance v2 / 1-5 / 1-0 (multi-modal cinematic), Kling 3.0 / 2-6, Google Veo 3-1, MiniMax Hailuo 2-3, ByteDance Dreamina 3-0. Covers text-to-video (t2v), image-to-video (i2v), and Veo's video-extend endpoint. The skill picks the right model for the user's intent (Arena-#1 qualitai-musicgenmedia-labs714KGenerate AI music on RunComfy via the `runcomfy` CLI — a smart router across the music-model catalog. Routes to ElevenLabs AI Music Generation (premium 44.1 kHz stereo vocal tracks, 5 s–5 min, $0.0083/s) and ACE Step / ACE Step 1.5 (StepFun-AI open-weights, tag-driven composition, multilingual lyrics, $0.0002–0.0003/s, ~27× cheaper), plus ACE Step audio-inpaint (regenerate a time range inside an existing track) and ACE Step audio-outpaint (extend a track before or after). Picks the right model fimage-to-videogenmedia-labs713KAnimate any still image on RunComfy — this skill is a smart router that matches the user's intent to the right i2v model in the RunComfy catalog. Picks HappyHorse 1.0 I2V (Arena #1, native audio, identity preservation) for general animations, Wan 2.7 with `audio_url` for custom-voiceover lip-sync, or Seedance 2.0 Pro for multi-modal animation from image + reference video + reference audio. Bundles each model's documented prompting patterns so the caller gets sharper output without burning iterat

Search skills and MCP servers

Fuzzy search across 23,137 skills and servers