Agent Skills

video-to-sfx

videosonilo-ai1.4K installs

Generate sound effects matched to a video using Sonilo — footsteps, impacts, ambience, foley — optionally scripted to specific timed segments, returning either the audio or a new video with the SFX muxed in. Use when the user has footage that needs sound design. For SFX from a text description alone, use the text-to-sfx skill; for music, use video-to-music.

Install

npx skills add https://github.com/sonilo-ai/skills --skill video-to-sfx
SKILL.md

Sonilo Video-to-SFX

Hand Sonilo a video and it generates sound effects matching what it sees — footsteps, impacts, ambience, UI sounds, whatever the scene calls for — or pin specific sounds to specific moments with segments. Generation runs as an async task on the backend; the tools poll internally and hand back the saved file.

Setup: See the setup-api-key skill.

⚠️ Cost: every tool below makes an API call that may incur charges. Only call it when explicitly requested.

Transport: MCP or CLI

Pick one at the start of the session and stay on it. Do not mix the two inside a single job, and do not announce the choice.

  1. Sonilo MCP tools visible in this session (video_to_sfx and friends) — use them. This is the preferred path: it needs no shell, and it is the only one that survives a very long generation. If a call fails to authenticate — rather than failing on its inputs — this transport is not usable in this session: go to 2 instead of retrying it.
  2. No usable Sonilo MCP tools, but sonilo account exits 0 — use the CLI commands below. Same API, same account, same credential file. Probe with sonilo account, not sonilo whoami: whoami exits 0 even when signed out, so it cannot tell the two states apart.
  3. Neither — stop and run the setup-api-key skill. Do not call api.sonilo.com with curl to work around it; both transports handle uploads, polling and retries that a bare request does not.

Quick Start

MCP tool call (recommended)

video_to_sfx(
    video_path="~/Desktop/action-scene.mp4",
    prompt="Footsteps on gravel, distant traffic, a door slam"
)

Python (pip install sonilo)

from sonilo import Sonilo

client = Sonilo()  # reads SONILO_API_KEY

foley = client.video_to_sfx.generate(video="action-scene.mp4", prompt="Footsteps on gravel, distant traffic, a door slam")
foley.save("foley.wav")

# video_to_video_sfx: get the video back with the effects muxed in
video = client.video_to_video_sfx.generate(video="action-scene.mp4", segments=[{"start": 0, "end": 2, "prompt": "footsteps on gravel"}])
video.save("with_sfx.mp4")

JavaScript / TypeScript (npm install sonilo)

import { SoniloClient } from "sonilo";

const client = new SoniloClient(); // reads SONILO_API_KEY

const foley = await client.videoToSfx.generate({
  video: "./action-scene.mp4",
  prompt: "Footsteps on gravel, distant traffic, a door slam",
});

// video_to_video_sfx: get the video back with the effects muxed in
const video = await client.videoToVideoSfx.generate({
  video: "./action-scene.mp4",
  segments: [{ start: 0, end: 2, prompt: "footsteps on gravel" }],
});

CLI (npm install -g sonilo-cli or pip install sonilo-cli)

sonilo video-to-sfx --video action-scene.mp4 --output foley.wav

Always async under the hood — the CLI submits and polls for you. --format accepts wav|mp3|aac|flac.

# the muxed video, from the CLI
sonilo video-to-video-sfx --video clip.mp4 --prompt "footsteps, distant thunder" --output foley.mp4

cURL (raw REST API, no MCP host)

curl -X POST "https://api.sonilo.com/v1/video-to-sfx" \
  -H "Authorization: Bearer $SONILO_API_KEY" \
  -F "video=@action-scene.mp4" \
  -F "prompt=Footsteps on gravel, distant traffic, a door slam"
# -> {"task_id": "..."}  poll GET /v1/tasks/{task_id} until status is succeeded/failed

Every call is task-based: the endpoint returns {"task_id": ...} (HTTP 202), and the result is fetched from GET /v1/tasks/{task_id} once status is terminal. The MCP tools do this polling for you and return the saved path directly — you only see the task_id if the call times out (see task-recovery).

Tools

Tool Description
video_to_sfx(video_path? | video_url?, prompt?, segments?, audio_format?, output_directory?) Generate SFX matched to a video. Returns audio only (not the source video).
video_to_video_sfx(video_path? | video_url?, prompt?, segments?, output_directory?) Same, but returns a new .mp4 with the SFX muxed in.

Parameters

Parameter Type Default Notes
prompt string — Optional overall description (max 2000 chars) — omit it to let Sonilo interpret the video on its own.
video_path string — .mp4/.mov/.webm/.m4v/.gif (gif must be animated) — a narrower set than the music tools. Max 480s (8 min), subject to the account's upload-size cap.
video_url string — HTTPS/HTTP URL to a video. Exactly one of video_path/video_url.
segments list[dict] — Script SFX to specific time ranges: [{"start": float, "end": float, "prompt": str}, ...]. See rules below. Max 30 segments.
audio_format string aac (.m4a) wav, mp3, aac, or flac. video_to_sfx only (video-to-video always outputs .mp4).
output_directory string SONILO_MCP_BASE_PATH Absolute, or relative to the base path.

segments rules

Validated by the backend before any charge — an invalid list is rejected with a 422/400 and nothing is billed:

  • First segment's start must be 0.
  • Segments must be contiguous: each end must equal the next segment's start.
  • Every end must be greater than its start.
  • Every prompt must be non-empty, max 200 chars.
  • The last end must not exceed the video's actual duration.
  • Max 30 segments total.

Prompting

No prompt is required — the model reads the cut. Quality comes from a time-segmented action map: what is on screen, what it's made of, what it does, second by second. The footage is the source of truth.

Before a paid call: probe the exact duration and existing audio, respect the 480 s cap (over = 422 reject, never truncated), and get sign-off — failed runs auto-refund, but your own retry is a new charge.

Workflow Tips

  • Leave prompt/segments unset to let Sonilo read the whole video and decide; use segments when you need specific sounds pinned to specific moments (e.g. a punch landing at 2.3s, a door slam at 5.0s).
  • Want the video back with SFX baked in? Use video_to_video_sfx instead of video_to_sfx.
  • Prompting: be specific and combine elements — "Heavy rain on a tin roof" beats "Rain".
  • Don't confuse this with music. For a background score or soundtrack, use video-to-music instead. To generate both music and SFX together in one balanced, single-charge call, use video-to-sound.
  • No footage? text-to-sfx generates a single clip from a description alone.
  • Don't know what it should sound like? Run video-analysis first: one call returns a sound-design brief read off the footage — shot-sized sfx_segments plus a ready-to-use whole-clip sfx_prompt (and a music brief too, unless you pass mode="sfx") — which beats guessing a prompt and rerolling. It is a paid call that generates nothing, so use it when the brief is genuinely unclear — not when the user already told you what they want.

Recovering a Timed-Out Call

Every tool here is async on the backend already; a long generation can still exceed TIME_OUT_SECONDS. If it does, the error carries a task_id — the job keeps running (and is already charged). Call get_sfx_task(task_id) — get_generation_task(task_id) on the hosted server — later to retrieve the result; see task-recovery.

Output Files

  • video_to_sfx: saved in the requested audio_format (.wav/.mp3/.flac, or .m4a for the aac default), named from the prompt (slugified) or sfx-<first 8 chars of the task id>.
  • video_to_video_sfx: a single .mp4 with the SFX muxed in.

Error Handling

Common errors: 401 invalid key, 402 insufficient balance / trial exhausted, 413 file too large, 422 invalid parameters or malformed segments, 429 rate limit. See the account skill.

Related skills

video-editgenmedia-labs715KEdit existing video on RunComfy — this skill is a smart router that matches the user's intent to the right edit model in the RunComfy catalog. Picks Wan 2.7 Edit-Video (general restyle / background swap / packaging swap, identity + motion preservation), Kling 2.6 Pro Motion Control (transfer precise motion from a reference video to a target character), or Lucy Edit Restyle (lightweight identity-stable restyle / outfit swap). Bundles each model's documented prompting patterns so the skill gets shai-video-generationgenmedia-labs714KGenerate AI videos on RunComfy via the `runcomfy` CLI — a smart router across the full video-model catalog: HappyHorse 1.0 (Arena #1, native in-pass audio), Wan-AI Wan 2-7 (open weights, audio-driven lip-sync), ByteDance Seedance v2 / 1-5 / 1-0 (multi-modal cinematic), Kling 3.0 / 2-6, Google Veo 3-1, MiniMax Hailuo 2-3, ByteDance Dreamina 3-0. Covers text-to-video (t2v), image-to-video (i2v), and Veo's video-extend endpoint. The skill picks the right model for the user's intent (Arena-#1 qualitai-musicgenmedia-labs714KGenerate AI music on RunComfy via the `runcomfy` CLI — a smart router across the music-model catalog. Routes to ElevenLabs AI Music Generation (premium 44.1 kHz stereo vocal tracks, 5 s–5 min, $0.0083/s) and ACE Step / ACE Step 1.5 (StepFun-AI open-weights, tag-driven composition, multilingual lyrics, $0.0002–0.0003/s, ~27× cheaper), plus ACE Step audio-inpaint (regenerate a time range inside an existing track) and ACE Step audio-outpaint (extend a track before or after). Picks the right model fimage-to-videogenmedia-labs713KAnimate any still image on RunComfy — this skill is a smart router that matches the user's intent to the right i2v model in the RunComfy catalog. Picks HappyHorse 1.0 I2V (Arena #1, native audio, identity preservation) for general animations, Wan 2.7 with `audio_url` for custom-voiceover lip-sync, or Seedance 2.0 Pro for multi-modal animation from image + reference video + reference audio. Bundles each model's documented prompting patterns so the caller gets sharper output without burning iterat

Search skills and MCP servers

Fuzzy search across 23,137 skills and servers