Agent Skills

listenhub

videomarswaveai1.5K installs

ListenHub CLI skills router. Routes to the correct skill based on user intent. Triggers on: "make a podcast", "explainer video", "read aloud", "TTS", "generate image", "generate video", "做播客", "解说视频", "朗读", "生成图片", "生成视频", "幻灯片", "slides", "音乐", "music", "generate music", "翻唱", "混音", "remix", "续写", "extend", "纯音乐", "instrumental", "配乐", "soundtrack", "分轨", "stem", "识别歌词", "克隆人声", "vocal clone", "cover song", "pixverse", "口型", "lipsync", "对口型", "parse URL", "解析链接", "提取内容".

Install

npx skills add https://github.com/marswaveai/skills --skill listenhub
SKILL.md

Purpose

This is a router skill. When users trigger a general ListenHub action, this skill identifies the intent and delegates to the appropriate specialized skill.

Routing Table

User intent Keywords Route to
ListenHub Voice end-to-end audio "端到端音频", "语音生成", "图片转音频", "图片生成音频", "多音色对白", "参考音频克隆", "克隆音色", "音效", "生成音效" /listenhub-voice
Podcast "podcast", "播客", "debate", "dialogue" /podcast
Explainer video "explainer", "解说视频", "tutorial video" /explainer
Slides / PPT "slides", "幻灯片", "PPT", "presentation" /slides
Voice cloning (persistent) "克隆我的声音", "克隆音色", "声音克隆", "语音克隆", "用我的声音", "自定义音色", "clone my voice", "custom voice" /voice-clone
TTS / Read aloud "TTS", "read aloud", "朗读", "配音", "语音合成" /tts
Image generation "generate image", "画一张", "生成图片", "AI图" /image-gen
Video generation "video", "视频", "seedance", "pixverse", "生成视频", "text to video", "做视频", "口型", "lipsync", "对口型" /video-gen
Music "music", "音乐", "生成音乐", "翻唱", "混音", "remix", "续写", "extend", "纯音乐", "instrumental", "配乐", "soundtrack", "分轨", "stem", "识别歌词", "克隆人声", "vocal clone" /music
Content extraction "parse URL", "extract content", "解析链接" /content-parser
Audio transcription "transcribe", "ASR", "语音转文字" /asr
Creator workflow "创作", "写公众号", "小红书", "口播" /creator

How to Route

  1. Read the user's message and identify which category it falls into
  2. Tell the user which skill you're routing to
  3. Follow that skill's SKILL.md completely

If the intent is ambiguous, ask the user to clarify:

Question: "What would you like to create?"
Options:
  - "ListenHub Voice" — End-to-end audio: sound effects, multi-voice dialogue, reference-audio cloning, image→audio
  - "Voice Cloning" — Clone your own voice into a reusable speaker
  - "Podcast" — Audio discussion on a topic
  - "Explainer Video" — Narrated video with AI visuals
  - "Slides" — Slide deck / presentation
  - "Music" — AI music: generate, remix, instrumental, soundtrack, stem, vocal clone
  - "Video" — AI video generation from text or reference materials

Prerequisites

Most skills require the ListenHub CLI. Check:

listenhub auth status --json

If not installed or not logged in, handle automatically — never ask the user to run install commands:

  1. Auto-install: run npm install -g @marswave/listenhub-cli silently
  2. Auto-login: run listenhub auth login (opens browser)

Exception: /asr runs locally and needs no CLI or API key.

Related skills

video-editgenmedia-labs715KEdit existing video on RunComfy — this skill is a smart router that matches the user's intent to the right edit model in the RunComfy catalog. Picks Wan 2.7 Edit-Video (general restyle / background swap / packaging swap, identity + motion preservation), Kling 2.6 Pro Motion Control (transfer precise motion from a reference video to a target character), or Lucy Edit Restyle (lightweight identity-stable restyle / outfit swap). Bundles each model's documented prompting patterns so the skill gets shai-video-generationgenmedia-labs714KGenerate AI videos on RunComfy via the `runcomfy` CLI — a smart router across the full video-model catalog: HappyHorse 1.0 (Arena #1, native in-pass audio), Wan-AI Wan 2-7 (open weights, audio-driven lip-sync), ByteDance Seedance v2 / 1-5 / 1-0 (multi-modal cinematic), Kling 3.0 / 2-6, Google Veo 3-1, MiniMax Hailuo 2-3, ByteDance Dreamina 3-0. Covers text-to-video (t2v), image-to-video (i2v), and Veo's video-extend endpoint. The skill picks the right model for the user's intent (Arena-#1 qualitai-musicgenmedia-labs714KGenerate AI music on RunComfy via the `runcomfy` CLI — a smart router across the music-model catalog. Routes to ElevenLabs AI Music Generation (premium 44.1 kHz stereo vocal tracks, 5 s–5 min, $0.0083/s) and ACE Step / ACE Step 1.5 (StepFun-AI open-weights, tag-driven composition, multilingual lyrics, $0.0002–0.0003/s, ~27× cheaper), plus ACE Step audio-inpaint (regenerate a time range inside an existing track) and ACE Step audio-outpaint (extend a track before or after). Picks the right model fimage-to-videogenmedia-labs713KAnimate any still image on RunComfy — this skill is a smart router that matches the user's intent to the right i2v model in the RunComfy catalog. Picks HappyHorse 1.0 I2V (Arena #1, native audio, identity preservation) for general animations, Wan 2.7 with `audio_url` for custom-voiceover lip-sync, or Seedance 2.0 Pro for multi-modal animation from image + reference video + reference audio. Bundles each model's documented prompting patterns so the caller gets sharper output without burning iterat

Search skills and MCP servers

Fuzzy search across 23,137 skills and servers