Agent Skills

elevenlabs-dubbing

ElevenLabs automatic dubbing - translate and dub audio/video into 29 languages while preserving speaker voice via inference.sh CLI. Capabilities: auto speaker detection, voice-preserving translation, video dubbing, audio localization. Use for: content localization, video translation, multilingual content, international distribution. Triggers: dubbing, dub video, translate audio, video translation, audio translation, localize content, elevenlabs dubbing, eleven labs dub, multilingual dub, voice t

Install

npx skills add https://github.com/inference-sh/skills --skill elevenlabs-dubbing
SKILL.md

Install the belt CLI skill: npx skills add belt-sh/cli

ElevenLabs Dubbing

Automatically dub audio and video into 29 languages via inference.sh CLI.

Dubbing

Quick Start

Requires inference.sh CLI (belt). Install instructions

belt login

# Dub English video to Spanish
belt app run elevenlabs/dubbing --input '{
  "audio": "https://video.mp4",
  "target_lang": "es"
}'

Supported Languages

Code Language Code Language
en English ko Korean
es Spanish ru Russian
fr French tr Turkish
de German nl Dutch
it Italian sv Swedish
pt Portuguese da Danish
pl Polish fi Finnish
hi Hindi no Norwegian
ar Arabic cs Czech
zh Chinese el Greek
ja Japanese he Hebrew
hu Hungarian id Indonesian
ms Malay ro Romanian
th Thai uk Ukrainian
vi Vietnamese

Supported Input Formats

  • MP3, MP4, WAV, MOV

Examples

Dub Video to Spanish

belt app run elevenlabs/dubbing --input '{
  "audio": "https://english-video.mp4",
  "target_lang": "es"
}'

Dub Audio to French

belt app run elevenlabs/dubbing --input '{
  "audio": "https://podcast-episode.mp3",
  "target_lang": "fr"
}'

Specify Source Language

# Skip auto-detection, specify source
belt app run elevenlabs/dubbing --input '{
  "audio": "https://german-video.mp4",
  "source_lang": "de",
  "target_lang": "en"
}'

Multi-Language Distribution

# Dub to multiple languages
for lang in es fr de ja ko; do
  belt app run elevenlabs/dubbing --input "{
    \"audio\": \"https://video.mp4\",
    \"target_lang\": \"$lang\"
  }" > "dubbed_${lang}.json"
  echo "Dubbed to $lang"
done

Features

  • Auto Speaker Detection: Identifies multiple speakers automatically
  • Voice Preservation: Maintains original speaker voice characteristics
  • Timing: Matches original speech timing and pacing
  • Multi-Speaker: Handles videos with multiple speakers

Workflow: Localize Content Pipeline

# 1. Start with original video
# 2. Dub to target language
belt app run elevenlabs/dubbing --input '{
  "audio": "https://original-video.mp4",
  "target_lang": "es"
}' > dubbed.json

# 3. Add subtitles in target language
belt app run elevenlabs/stt --input '{
  "audio": "<dubbed-audio-url>",
  "language_code": "spa"
}' > transcript.json

# 4. Caption the dubbed video
belt app run infsh/caption-videos --input '{
  "video_file": "<dubbed-video-url>",
  "segments": [{"start": 0.0, "end": 2.5, "text": "<text-from-transcript>"}]
}'

Use Cases

  • Content Creators: Reach international audiences
  • E-learning: Localize courses for global students
  • Marketing: Adapt campaigns for different markets
  • Podcasts: Distribute in multiple languages
  • Corporate: Multilingual training and communications
  • Film/TV: Quick dubbing for distribution

Related Skills

# ElevenLabs TTS (generate speech in any language)
npx skills add inference-sh/skills@elevenlabs-tts

# ElevenLabs STT (transcribe dubbed content)
npx skills add inference-sh/skills@elevenlabs-stt

# ElevenLabs voice changer (transform voices)
npx skills add inference-sh/skills@elevenlabs-voice-changer

# Full platform skill (all apps)
npx skills add inference-sh/skills@infsh-cli

Browse all audio apps: belt app list --category audio

Related skills

video-editgenmedia-labs715KEdit existing video on RunComfy — this skill is a smart router that matches the user's intent to the right edit model in the RunComfy catalog. Picks Wan 2.7 Edit-Video (general restyle / background swap / packaging swap, identity + motion preservation), Kling 2.6 Pro Motion Control (transfer precise motion from a reference video to a target character), or Lucy Edit Restyle (lightweight identity-stable restyle / outfit swap). Bundles each model's documented prompting patterns so the skill gets shai-video-generationgenmedia-labs714KGenerate AI videos on RunComfy via the `runcomfy` CLI — a smart router across the full video-model catalog: HappyHorse 1.0 (Arena #1, native in-pass audio), Wan-AI Wan 2-7 (open weights, audio-driven lip-sync), ByteDance Seedance v2 / 1-5 / 1-0 (multi-modal cinematic), Kling 3.0 / 2-6, Google Veo 3-1, MiniMax Hailuo 2-3, ByteDance Dreamina 3-0. Covers text-to-video (t2v), image-to-video (i2v), and Veo's video-extend endpoint. The skill picks the right model for the user's intent (Arena-#1 qualitai-musicgenmedia-labs714KGenerate AI music on RunComfy via the `runcomfy` CLI — a smart router across the music-model catalog. Routes to ElevenLabs AI Music Generation (premium 44.1 kHz stereo vocal tracks, 5 s–5 min, $0.0083/s) and ACE Step / ACE Step 1.5 (StepFun-AI open-weights, tag-driven composition, multilingual lyrics, $0.0002–0.0003/s, ~27× cheaper), plus ACE Step audio-inpaint (regenerate a time range inside an existing track) and ACE Step audio-outpaint (extend a track before or after). Picks the right model fimage-to-videogenmedia-labs713KAnimate any still image on RunComfy — this skill is a smart router that matches the user's intent to the right i2v model in the RunComfy catalog. Picks HappyHorse 1.0 I2V (Arena #1, native audio, identity preservation) for general animations, Wan 2.7 with `audio_url` for custom-voiceover lip-sync, or Seedance 2.0 Pro for multi-modal animation from image + reference video + reference audio. Bundles each model's documented prompting patterns so the caller gets sharper output without burning iterat

Search skills and MCP servers

Fuzzy search across 23,137 skills and servers