Agent Skills

elevenlabs-voice-isolator

ElevenLabs voice isolator - remove background noise and isolate vocals from audio via inference.sh CLI. Capabilities: noise removal, voice extraction, audio cleanup, background removal. Use for: podcast cleanup, interview audio, music vocals, noisy recordings, audio restoration. Triggers: voice isolator, noise removal, background removal, isolate voice, clean audio, remove background noise, audio cleanup, voice extraction, elevenlabs isolator, eleven labs noise, vocal isolation, denoise, audio r

Install

npx skills add https://github.com/inference-sh/skills --skill elevenlabs-voice-isolator
SKILL.md

Install the belt CLI skill: npx skills add belt-sh/cli

ElevenLabs Voice Isolator

Remove background noise and isolate voices from audio via inference.sh CLI.

Voice Isolator

Quick Start

Requires inference.sh CLI (belt). Install instructions

belt login

# Isolate voice from noisy audio
belt app run elevenlabs/voice-isolator --input '{"audio": "https://noisy-recording.mp3"}'

Supported Formats

Format Max Size Max Duration
WAV 500MB 1 hour
MP3 500MB 1 hour
FLAC 500MB 1 hour
OGG 500MB 1 hour
AAC 500MB 1 hour

Examples

Clean Up a Recording

# Remove background noise from a podcast recording
belt app run elevenlabs/voice-isolator --input '{"audio": "https://noisy-podcast.mp3"}'

Clean Interview Audio

# Isolate speaker from café background noise
belt app run elevenlabs/voice-isolator --input '{"audio": "https://cafe-interview.mp3"}'

Extract Vocals from Music

# Separate vocals from instrumental
belt app run elevenlabs/voice-isolator --input '{"audio": "https://song.mp3"}'

What It Removes

  • Ambient/environmental noise
  • Background music
  • Reverb and echo
  • Wind noise
  • Traffic and crowd noise
  • Electrical hum/buzz
  • Other non-voice sounds

Workflow: Clean → Transcribe

# 1. Isolate voice from noisy recording
belt app run elevenlabs/voice-isolator --input '{
  "audio": "https://noisy-meeting.mp3"
}' > cleaned.json

# 2. Transcribe the clean audio
belt app run elevenlabs/stt --input '{
  "audio": "<cleaned-audio-url>",
  "diarize": true
}'

Workflow: Clean → Voice Change

# 1. Clean up the audio
belt app run elevenlabs/voice-isolator --input '{
  "audio": "https://raw-recording.mp3"
}' > cleaned.json

# 2. Transform the voice
belt app run elevenlabs/voice-changer --input '{
  "audio": "<cleaned-audio-url>",
  "voice": "george"
}'

Workflow: Clean → Add to Video

# 1. Clean the voiceover
belt app run elevenlabs/voice-isolator --input '{
  "audio": "https://raw-voiceover.mp3"
}' > cleaned.json

# 2. Merge with video
belt app run infsh/video-audio-merger --input '{
  "video_file": "video.mp4",
  "audio_file": "<cleaned-audio-url>"
}'

Use Cases

  • Podcasts: Clean up recordings with background noise
  • Interviews: Remove café/office ambient sounds
  • Music: Extract vocals for remixes or karaoke
  • Video Production: Clean dialogue audio
  • Archival: Restore old or degraded recordings
  • Meetings: Improve recording clarity
  • Voice Cloning Prep: Clean source audio for better cloning results

Related Skills

# ElevenLabs voice changer (transform voice after cleaning)
npx skills add inference-sh/skills@elevenlabs-voice-changer

# ElevenLabs STT (transcribe clean audio)
npx skills add inference-sh/skills@elevenlabs-stt

# ElevenLabs TTS (generate clean speech from text)
npx skills add inference-sh/skills@elevenlabs-tts

# Full platform skill (all apps)
npx skills add inference-sh/skills@infsh-cli

Browse all audio apps: belt app list --category audio

Related skills

video-editgenmedia-labs715KEdit existing video on RunComfy — this skill is a smart router that matches the user's intent to the right edit model in the RunComfy catalog. Picks Wan 2.7 Edit-Video (general restyle / background swap / packaging swap, identity + motion preservation), Kling 2.6 Pro Motion Control (transfer precise motion from a reference video to a target character), or Lucy Edit Restyle (lightweight identity-stable restyle / outfit swap). Bundles each model's documented prompting patterns so the skill gets shai-video-generationgenmedia-labs714KGenerate AI videos on RunComfy via the `runcomfy` CLI — a smart router across the full video-model catalog: HappyHorse 1.0 (Arena #1, native in-pass audio), Wan-AI Wan 2-7 (open weights, audio-driven lip-sync), ByteDance Seedance v2 / 1-5 / 1-0 (multi-modal cinematic), Kling 3.0 / 2-6, Google Veo 3-1, MiniMax Hailuo 2-3, ByteDance Dreamina 3-0. Covers text-to-video (t2v), image-to-video (i2v), and Veo's video-extend endpoint. The skill picks the right model for the user's intent (Arena-#1 qualitai-musicgenmedia-labs714KGenerate AI music on RunComfy via the `runcomfy` CLI — a smart router across the music-model catalog. Routes to ElevenLabs AI Music Generation (premium 44.1 kHz stereo vocal tracks, 5 s–5 min, $0.0083/s) and ACE Step / ACE Step 1.5 (StepFun-AI open-weights, tag-driven composition, multilingual lyrics, $0.0002–0.0003/s, ~27× cheaper), plus ACE Step audio-inpaint (regenerate a time range inside an existing track) and ACE Step audio-outpaint (extend a track before or after). Picks the right model fimage-to-videogenmedia-labs713KAnimate any still image on RunComfy — this skill is a smart router that matches the user's intent to the right i2v model in the RunComfy catalog. Picks HappyHorse 1.0 I2V (Arena #1, native audio, identity preservation) for general animations, Wan 2.7 with `audio_url` for custom-voiceover lip-sync, or Seedance 2.0 Pro for multi-modal animation from image + reference video + reference audio. Bundles each model's documented prompting patterns so the caller gets sharper output without burning iterat

Search skills and MCP servers

Fuzzy search across 23,137 skills and servers