Agent Skills

fish-audio

videoacedatacloud4.1K installs

Generate AI text-to-speech audio, use saved voices, or create a one-shot voice clone from an HTTPS reference audio URL and exact transcript via AceDataCloud API.

Install

npx skills add https://github.com/acedatacloud/skills --skill fish-audio
SKILL.md

Fish Audio — Text-to-Speech

Generate narration / voiceover through AceDataCloud's Fish Audio API.

Setup: See authentication for token setup.

Quick Start

curl -X POST https://api.acedata.cloud/fish/tts \
  -H "Authorization: ******ACEDATACLOUD_API_TOKEN" \
  -H "Content-Type: application/json" \
  -H "model: s2-pro" \
  -d '{"text":"你好,欢迎使用 AceData Cloud。","reference_id":"d7900c21663f485ab63ebdb7e5905036","format":"mp3"}'

Synchronous responses return a direct audio URL:

{"audio_url":"https://platform.r2.fish.audio/task/8a72ff9840234006a9f74cb2fa04f978.mp3"}

Endpoints

Endpoint Purpose
POST /fish/tts Text-to-speech generation
GET /fish/model Browse/search public Fish reference voices
GET /fish/model/{id} Fetch one reference voice by ID
POST /fish/tasks Poll async TTS jobs when async: true

Workflows

1. Find a reference voice

curl "https://api.acedata.cloud/fish/model?page_size=10&page_number=1&title=Marcus" \
  -H "Authorization: ******ACEDATACLOUD_API_TOKEN"

The response includes items[] with public voice metadata such as _id, title, languages, tags, visibility, and state. Use an item _id as reference_id in TTS requests.

2. Text-to-Speech

POST /fish/tts
Headers:
  model: s2-pro

{
  "text": "Your narration text.",
  "reference_id": "d7900c21663f485ab63ebdb7e5905036",
  "format": "mp3"
}

3. One-shot voice cloning

Use a temporary reference voice without creating a persistent model:

POST /fish/tts
Headers:
  model: s2-pro

{
  "text": "New speech in the referenced voice.",
  "format": "mp3",
  "references": [{
    "audio": "https://cdn.acedata.cloud/reference.mp3",
    "text": "The exact words spoken in the reference audio."
  }]
}

audio must be a public HTTPS MP3/WAV URL and text must be the exact transcript. Use one reference lasting 10–270 seconds. Do not combine references with reference_id; use reference_id when the same saved/public voice will be reused. Raw bytes, Base64, data URIs, and MessagePack are not accepted by the AceDataCloud endpoint.

4. Async TTS

POST /fish/tts
Headers:
  model: s1

{
  "text": "Longer narration for background processing.",
  "async": true,
  "callback_url": "https://api.acedata.cloud/health"
}

Async: See async task polling. Poll via POST /fish/tasks with {"id":"..."}.

Parameters — /fish/tts

Header

Parameter Values Description
model "s1", "s2-pro", "s2.1-pro" Fish TTS engine selection

JSON body

Parameter Type / Values Description
text string Text to synthesize (required)
reference_id string Public/reference voice ID from GET /fish/model
format "mp3", "wav", "pcm" Output format
sample_rate integer Optional output sample rate
mp3_bitrate 64, 128, 192 MP3 bitrate
latency "normal", "balanced" TTS latency mode
chunk_length / min_chunk_length integer Chunking controls
temperature, top_p, repetition_penalty number Sampling controls
max_new_tokens integer Maximum generated tokens
normalize boolean Normalize generated audio
prosody object Prosody tuning
references array One {audio, text} object for a one-shot voice clone; mutually exclusive with reference_id
callback_url string Async callback URL
async boolean Run asynchronously and poll /fish/tasks

Gotchas

  • The documented TTS endpoint is POST /fish/tts — not /fish/audios.
  • Choose the Fish engine with the model request header, not a JSON model field.
  • Use reference_id from GET /fish/model — not voice_id.
  • Use references for a one-shot clone that is not saved as a model.
  • Billing is based on the target text UTF-8 byte count; the reference audio does not add a separate clone fee.
  • Synchronous requests return audio_url directly; async jobs should be polled via /fish/tasks.

Related skills

video-editgenmedia-labs715KEdit existing video on RunComfy — this skill is a smart router that matches the user's intent to the right edit model in the RunComfy catalog. Picks Wan 2.7 Edit-Video (general restyle / background swap / packaging swap, identity + motion preservation), Kling 2.6 Pro Motion Control (transfer precise motion from a reference video to a target character), or Lucy Edit Restyle (lightweight identity-stable restyle / outfit swap). Bundles each model's documented prompting patterns so the skill gets shai-video-generationgenmedia-labs714KGenerate AI videos on RunComfy via the `runcomfy` CLI — a smart router across the full video-model catalog: HappyHorse 1.0 (Arena #1, native in-pass audio), Wan-AI Wan 2-7 (open weights, audio-driven lip-sync), ByteDance Seedance v2 / 1-5 / 1-0 (multi-modal cinematic), Kling 3.0 / 2-6, Google Veo 3-1, MiniMax Hailuo 2-3, ByteDance Dreamina 3-0. Covers text-to-video (t2v), image-to-video (i2v), and Veo's video-extend endpoint. The skill picks the right model for the user's intent (Arena-#1 qualitai-musicgenmedia-labs714KGenerate AI music on RunComfy via the `runcomfy` CLI — a smart router across the music-model catalog. Routes to ElevenLabs AI Music Generation (premium 44.1 kHz stereo vocal tracks, 5 s–5 min, $0.0083/s) and ACE Step / ACE Step 1.5 (StepFun-AI open-weights, tag-driven composition, multilingual lyrics, $0.0002–0.0003/s, ~27× cheaper), plus ACE Step audio-inpaint (regenerate a time range inside an existing track) and ACE Step audio-outpaint (extend a track before or after). Picks the right model fimage-to-videogenmedia-labs713KAnimate any still image on RunComfy — this skill is a smart router that matches the user's intent to the right i2v model in the RunComfy catalog. Picks HappyHorse 1.0 I2V (Arena #1, native audio, identity preservation) for general animations, Wan 2.7 with `audio_url` for custom-voiceover lip-sync, or Seedance 2.0 Pro for multi-modal animation from image + reference video + reference audio. Bundles each model's documented prompting patterns so the caller gets sharper output without burning iterat

Search skills and MCP servers

Fuzzy search across 23,137 skills and servers