Agent Skills

audio-generation

Plan TTS, voice cloning, voice change, translated dub, or lip-sync audio. Resolve voice and reference policy before handing a ready request to voice-batch-runner.

Install

npx skills add https://github.com/postplusai/postplus-skills --skill audio-generation
SKILL.md

Audio Generation

Use When

  • The desired final asset is generated audio or audio prepared for a video render.
  • The request includes TTS, voice design, voice cloning, voice change, translated dub, podcast audio, or lip-sync handoff.
  • The next decision is audio task class, reference policy, and runner handoff.

Do Not Use When

  • The user needs speech-to-text from existing audio. Use audio-transcription; use its local subtitle reference when an existing timed transcript needs subtitle files.
  • The voice request is already normalized for execution. Use voice-batch-runner.
  • The final work is a full video production pipeline. Use video-batch-runner after the audio handoff is clear.

Core Boundary

This is the audio generation controller. It does not submit jobs.

It must classify the task and hand off execution. It must not let a runner invent voice strategy, translation policy, or lip-sync intent.

Task Classes

Task class Use when Handoff
tts new spoken audio from script voice-batch-runner with voice design rules
change_voice preserve script, alter voice identity or delivery reference contract, then voice-batch-runner
translate_dub translate and dub source audio require language, meaning-preservation, and timing policy
voice_clone_take approved reference voice should preserve timbre bind reference audio, then voice-batch-runner
podcast_audio speaker-led or conversational audio create voice/script handoff before video assembly
lip_sync_handoff audio drives talking-head or UGC render voice-batch-runner, then video-batch-runner

Reference Rules

  • Approved voice reference audio is binding.
  • Accent, energy, cadence, or genre examples are inspiration-only unless the user explicitly binds them.
  • Source audio used only for translation meaning is not a voice identity binding unless stated.
  • Excluded voices, music, or effects must not enter the runner request.

Routing Table

If not audio-generation Send to
Transcribe existing audio audio-transcription
Need generated image/video around audio video-batch-runner
Need normalized hosted voice execution voice-batch-runner
Need lip-sync video after audio video-batch-runner

Output Shape

Return:

  • taskClass
  • scriptPolicy
  • voicePolicy
  • referencePolicy
  • runnerHandoff
  • nextVideoHandoff when lip-sync or video assembly follows
  • mustNotDo

Stop Conditions

  • Stop when required user intent, source evidence, or owned input artifacts are missing and guessing would change the result.
  • Do not ask voice-batch-runner to decide the creative role of the voice.

Public Command Boundary

  • Choose the smallest matching command or workflow from the user input and run it directly.

  • This public skill is instruction-driven. Produce the controller handoff artifact directly from the available evidence.

  • Do not call private provider/runtime paths or unpublished local tools.

  • If the CLI returns a quote-confirmation challenge, obtain user approval for its scope and cost before running postplus quote confirm --json --challenge-file <challenge.json> and retry with the returned token.

Related skills

video-editgenmedia-labs715KEdit existing video on RunComfy — this skill is a smart router that matches the user's intent to the right edit model in the RunComfy catalog. Picks Wan 2.7 Edit-Video (general restyle / background swap / packaging swap, identity + motion preservation), Kling 2.6 Pro Motion Control (transfer precise motion from a reference video to a target character), or Lucy Edit Restyle (lightweight identity-stable restyle / outfit swap). Bundles each model's documented prompting patterns so the skill gets shai-video-generationgenmedia-labs714KGenerate AI videos on RunComfy via the `runcomfy` CLI — a smart router across the full video-model catalog: HappyHorse 1.0 (Arena #1, native in-pass audio), Wan-AI Wan 2-7 (open weights, audio-driven lip-sync), ByteDance Seedance v2 / 1-5 / 1-0 (multi-modal cinematic), Kling 3.0 / 2-6, Google Veo 3-1, MiniMax Hailuo 2-3, ByteDance Dreamina 3-0. Covers text-to-video (t2v), image-to-video (i2v), and Veo's video-extend endpoint. The skill picks the right model for the user's intent (Arena-#1 qualitai-musicgenmedia-labs714KGenerate AI music on RunComfy via the `runcomfy` CLI — a smart router across the music-model catalog. Routes to ElevenLabs AI Music Generation (premium 44.1 kHz stereo vocal tracks, 5 s–5 min, $0.0083/s) and ACE Step / ACE Step 1.5 (StepFun-AI open-weights, tag-driven composition, multilingual lyrics, $0.0002–0.0003/s, ~27× cheaper), plus ACE Step audio-inpaint (regenerate a time range inside an existing track) and ACE Step audio-outpaint (extend a track before or after). Picks the right model fimage-to-videogenmedia-labs713KAnimate any still image on RunComfy — this skill is a smart router that matches the user's intent to the right i2v model in the RunComfy catalog. Picks HappyHorse 1.0 I2V (Arena #1, native audio, identity preservation) for general animations, Wan 2.7 with `audio_url` for custom-voiceover lip-sync, or Seedance 2.0 Pro for multi-modal animation from image + reference video + reference audio. Bundles each model's documented prompting patterns so the caller gets sharper output without burning iterat

Search skills and MCP servers

Fuzzy search across 23,137 skills and servers