Agent Skills

text-to-sfx

videosonilo-ai1.4K installs

Generate a sound effect from a text description using Sonilo — a UI chime, a whoosh, an impact, ambience, a stylized cue — when there is no video to match. Use when the user describes the sound they want in words, with or without a length. For SFX matched to footage, use the video-to-sfx skill; for music, use text-to-music.

Install

npx skills add https://github.com/sonilo-ai/skills --skill text-to-sfx
SKILL.md

Sonilo Text-to-SFX

Generate a single sound effect from a text description — no video involved. The prompt IS the input, so describe the action and materials directly. Generation runs as an async task on the backend; the tool polls internally and hands back the saved file.

Setup: See the setup-api-key skill.

⚠️ Cost: this tool makes an API call that may incur charges. Only call it when explicitly requested.

Matching sound to footage instead? Use video-to-sfx — it reads the cut and can pin sounds to specific moments, which a text prompt cannot do.

Transport: MCP or CLI

Pick one at the start of the session and stay on it. Do not mix the two inside a single job, and do not announce the choice.

  1. Sonilo MCP tools visible in this session (text_to_sfx and friends) — use them. This is the preferred path: it needs no shell, and it is the only one that survives a very long generation. If a call fails to authenticate — rather than failing on its inputs — this transport is not usable in this session: go to 2 instead of retrying it.
  2. No usable Sonilo MCP tools, but sonilo account exits 0 — use the CLI commands below. Same API, same account, same credential file. Probe with sonilo account, not sonilo whoami: whoami exits 0 even when signed out, so it cannot tell the two states apart.
  3. Neither — stop and run the setup-api-key skill. Do not call api.sonilo.com with curl to work around it; both transports handle uploads, polling and retries that a bare request does not.

Quick Start

MCP tool call (recommended)

text_to_sfx(
    prompt="Thunder rumbling in the distance with light rain",
    duration=6
)

Python (pip install sonilo)

from sonilo import Sonilo

client = Sonilo()  # reads SONILO_API_KEY

sfx = client.text_to_sfx.generate(prompt="Thunder rumbling in the distance with light rain", duration=6)
sfx.save("thunder.m4a")

JavaScript / TypeScript (npm install sonilo)

import { SoniloClient } from "sonilo";

const client = new SoniloClient(); // reads SONILO_API_KEY

const sfx = await client.textToSfx.generate({
  prompt: "Thunder rumbling in the distance with light rain",
  duration: 6,
});

CLI (npm install -g sonilo-cli or pip install sonilo-cli)

sonilo text-to-sfx --prompt "Thunder rumbling in the distance with light rain" --duration 6

Always async under the hood — the CLI submits and polls for you. --format accepts wav|mp3|aac|flac.

cURL (raw REST API, no MCP host)

curl -X POST "https://api.sonilo.com/v1/text-to-sfx" \
  -H "Authorization: Bearer $SONILO_API_KEY" \
  --data-urlencode "prompt=Thunder rumbling in the distance with light rain" \
  --data-urlencode "duration=6"
# -> {"task_id": "..."}  poll GET /v1/tasks/{task_id} until status is succeeded/failed

The endpoint returns {"task_id": ...} (HTTP 202) and the result is fetched from GET /v1/tasks/{task_id} once status is terminal. The MCP tool does this polling for you and returns the saved path directly — you only see the task_id if the call times out (see task-recovery).

Tool

Tool Description
text_to_sfx(prompt, duration?, audio_format?, output_directory?) Generate one SFX clip from a text description only.

Parameters

Parameter Type Default Notes
prompt string — Required. 1–2000 chars.
duration number Sonilo's default, 8 s Optional, 0.5–180 seconds. Omit it when the user names no length — do not invent one. Fractional values are allowed: the shortest effects, a latch or a click, run well under a second.
audio_format string aac (.m4a) wav, mp3, aac, or flac.
output_directory string SONILO_MCP_BASE_PATH Absolute, or relative to the base path.

Prompting

The prompt is the only input — there is no footage to map. Describe the action and the materials directly, and combine elements: "Heavy rain on a tin roof" beats "Rain"; "Cinematic braam, horror" or "8-bit retro jump sound" for stylized cues. The same materials vocabulary and sound-bundle thinking as the video path applies: references/sfx-prompting.md.

Generate once and iterate on the prompt, not on rerolls — failed runs auto-refund, but your own retry is a new charge.

Workflow Tips

  • This is for a single clip with no video context — a UI chime, a whoosh, a foley element you'll layer yourself. If the user has footage, use video-to-sfx instead.
  • Duration is optional; never invent one. Pass the length the user named. If they named none, omit duration and Sonilo uses its default (8 s) — ask only when the length actually matters to the user.
  • Don't confuse this with music. For a background score or soundtrack, use text-to-music or video-to-music.

Recovering a Timed-Out Call

Async on the backend already; a long generation can still exceed TIME_OUT_SECONDS. If it does, the error carries a task_id — the job keeps running (and is already charged). Call get_sfx_task(task_id) — get_generation_task(task_id) on the hosted server — later to retrieve the result; see task-recovery.

Output Files

Saved in the requested audio_format (.wav/.mp3/.flac, or .m4a for the aac default), named from the prompt (slugified) or sfx-<first 8 chars of the task id>.

Error Handling

Common errors: 401 invalid key, 402 insufficient balance / trial exhausted, 422 invalid parameters (e.g. duration out of range), 429 rate limit. See the account skill.

Related skills

video-editgenmedia-labs715KEdit existing video on RunComfy — this skill is a smart router that matches the user's intent to the right edit model in the RunComfy catalog. Picks Wan 2.7 Edit-Video (general restyle / background swap / packaging swap, identity + motion preservation), Kling 2.6 Pro Motion Control (transfer precise motion from a reference video to a target character), or Lucy Edit Restyle (lightweight identity-stable restyle / outfit swap). Bundles each model's documented prompting patterns so the skill gets shai-video-generationgenmedia-labs714KGenerate AI videos on RunComfy via the `runcomfy` CLI — a smart router across the full video-model catalog: HappyHorse 1.0 (Arena #1, native in-pass audio), Wan-AI Wan 2-7 (open weights, audio-driven lip-sync), ByteDance Seedance v2 / 1-5 / 1-0 (multi-modal cinematic), Kling 3.0 / 2-6, Google Veo 3-1, MiniMax Hailuo 2-3, ByteDance Dreamina 3-0. Covers text-to-video (t2v), image-to-video (i2v), and Veo's video-extend endpoint. The skill picks the right model for the user's intent (Arena-#1 qualitai-musicgenmedia-labs714KGenerate AI music on RunComfy via the `runcomfy` CLI — a smart router across the music-model catalog. Routes to ElevenLabs AI Music Generation (premium 44.1 kHz stereo vocal tracks, 5 s–5 min, $0.0083/s) and ACE Step / ACE Step 1.5 (StepFun-AI open-weights, tag-driven composition, multilingual lyrics, $0.0002–0.0003/s, ~27× cheaper), plus ACE Step audio-inpaint (regenerate a time range inside an existing track) and ACE Step audio-outpaint (extend a track before or after). Picks the right model fimage-to-videogenmedia-labs713KAnimate any still image on RunComfy — this skill is a smart router that matches the user's intent to the right i2v model in the RunComfy catalog. Picks HappyHorse 1.0 I2V (Arena #1, native audio, identity preservation) for general animations, Wan 2.7 with `audio_url` for custom-voiceover lip-sync, or Seedance 2.0 Pro for multi-modal animation from image + reference video + reference audio. Bundles each model's documented prompting patterns so the caller gets sharper output without burning iterat

Search skills and MCP servers

Fuzzy search across 23,137 skills and servers