Agent Skills

audio-transcription

videopostplusai1.3K installs

Transcribe local or remote audio into text and timestamps. Convert an existing timed transcript locally into SRT or ASS without another transcription job.

Install

npx skills add https://github.com/postplusai/postplus-skills --skill audio-transcription
SKILL.md

Audio Transcription

Use When

  • The input is audio and the main job is speech-to-text, subtitle-ready timing, rough speech search, multilingual transcription, or durable transcript artifacts.
  • Use video-transcription for video inputs and media-analysis for semantic video understanding.

If a timed transcript already exists and only subtitle output is requested, skip transcription and read the local subtitle conversion reference.

Do Not Use When

  • The task needs new creative generation or visual analysis rather than speech or subtitles.
  • Required inputs are missing and guessing would change the result.

Execution Boundary

  • Hosted transcription runs through the public postplus media transcribe verb and is async. A submit records the run handle, current status, and completed artifacts when available.
  • Pass a local path, HTTPS URL, existing PostPlus media reference, or data URI directly to --audio. The CLI validates and prepares local media before the single hosted submit.
  • A higher-quality default model and a faster, cheaper variant are available; prefer the default when subtitle quality matters and use the cheaper variant for an explicit rough pass. The generated example below shows the default endpoint key.

Source And Path

  • Supply the media duration so PostPlus can validate the request before it runs; a missing duration fails before submission.
  • Request timestamps when the output will feed subtitles or edit decisions.
  • Start with one source file or audio URL before larger batches.
  • Keep internal requests, responses, manifests, normalized transcripts, and downloaded artifacts under .postplus/audio-transcription; keep final user-facing transcript exports outside .postplus.

Handoff

  • If status is pending, preserve the result path and follow the CLI-returned action or resume command for the same operation. Do not submit another job. Stop and report when the CLI wait/recovery boundary is reached.
  • When SRT/ASS is requested, use the actual timed transcript and read local subtitle conversion. Convert locally without another hosted request; do not invent a CLI export command.

Stop Conditions

  • Stop when required user intent, source evidence, or owned input artifacts are missing and guessing would change the result.

Public Command Boundary

  • Choose the smallest matching command or workflow from the user input and run it directly.

  • Readiness diagnostics: postplus doctor --skill audio-transcription.

  • Use postplus media schema --json only when you need the full endpoint, flag, and enum contract or are repairing an unknown request shape.

  • Run the hosted transcription job with the generated command below; do not use another execution interface.

  • Pass the source directly through --audio; do not pre-upload it or construct a manual request object.

postplus media transcribe transcription \
  --audio ./reference.wav \
  --duration-seconds 1 \
  --wait \
  --output ./result.json

Follow the CLI's structured result and reported next action; do not infer recovery from free-text messages. Wait for explicit user approval when requested; an action does not authorize spending, publishing, or overwriting. Resume the same operation through its returned checkpoint or action; never resubmit uncertain work, repeat exhausted recovery, or switch providers to bypass failure.

  • If the CLI returns a quote-confirmation challenge, obtain user approval for its scope and cost before running postplus quote confirm --json --challenge-file <challenge.json> and retry with the returned token.

Related skills

video-editgenmedia-labs715KEdit existing video on RunComfy — this skill is a smart router that matches the user's intent to the right edit model in the RunComfy catalog. Picks Wan 2.7 Edit-Video (general restyle / background swap / packaging swap, identity + motion preservation), Kling 2.6 Pro Motion Control (transfer precise motion from a reference video to a target character), or Lucy Edit Restyle (lightweight identity-stable restyle / outfit swap). Bundles each model's documented prompting patterns so the skill gets shai-video-generationgenmedia-labs714KGenerate AI videos on RunComfy via the `runcomfy` CLI — a smart router across the full video-model catalog: HappyHorse 1.0 (Arena #1, native in-pass audio), Wan-AI Wan 2-7 (open weights, audio-driven lip-sync), ByteDance Seedance v2 / 1-5 / 1-0 (multi-modal cinematic), Kling 3.0 / 2-6, Google Veo 3-1, MiniMax Hailuo 2-3, ByteDance Dreamina 3-0. Covers text-to-video (t2v), image-to-video (i2v), and Veo's video-extend endpoint. The skill picks the right model for the user's intent (Arena-#1 qualitai-musicgenmedia-labs714KGenerate AI music on RunComfy via the `runcomfy` CLI — a smart router across the music-model catalog. Routes to ElevenLabs AI Music Generation (premium 44.1 kHz stereo vocal tracks, 5 s–5 min, $0.0083/s) and ACE Step / ACE Step 1.5 (StepFun-AI open-weights, tag-driven composition, multilingual lyrics, $0.0002–0.0003/s, ~27× cheaper), plus ACE Step audio-inpaint (regenerate a time range inside an existing track) and ACE Step audio-outpaint (extend a track before or after). Picks the right model fimage-to-videogenmedia-labs713KAnimate any still image on RunComfy — this skill is a smart router that matches the user's intent to the right i2v model in the RunComfy catalog. Picks HappyHorse 1.0 I2V (Arena #1, native audio, identity preservation) for general animations, Wan 2.7 with `audio_url` for custom-voiceover lip-sync, or Seedance 2.0 Pro for multi-modal animation from image + reference video + reference audio. Bundles each model's documented prompting patterns so the caller gets sharper output without burning iterat

Search skills and MCP servers

Fuzzy search across 23,137 skills and servers