Agent Skills

Use when an AI agent needs to work with VidMuse projects, account state, subscription credits, video/audio/image/text/tool models, generated or uploaded assets, style library entries, voice library entries, generation parameters, threads, messages, or long-term memory. Installs and uses the `vidmuse` command runtime when needed.

Install

npx skills add https://github.com/sandai-org/vidmuse-skills --skill vidmuse
SKILL.md

VidMuse Skill

Use this skill when a task asks for VidMuse account state, subscription plan, model discovery, assets, style library entries, voice library entries, generation params, thread/message operations, or long-term memory operations. Treat the vidmuse command as the runtime for accessing VidMuse capabilities; keep user-facing answers focused on the VidMuse task unless the user asks about the command itself.

Workflow

  1. Ensure the vidmuse command runtime exists.

    macOS, Linux, WSL, or Git Bash:

    if vidmuse_path="$(command -v vidmuse)"; then
      :
    else
      tmp="$(mktemp)"
      curl -fsSL "https://vidmuse.sandcdn.com/cli/install.sh" -o "$tmp"
      installed_path="$(VIDMUSE_CLI_VERSION="${VIDMUSE_CLI_VERSION:-latest}" bash "$tmp" | tail -n 1)"
      installed_dir="$(dirname "$installed_path")"
      case ":${PATH:-}:" in *":${installed_dir}:"*) ;; *) export PATH="${installed_dir}:${PATH:-}" ;; esac
    fi
    vidmuse --version
    

    Windows PowerShell:

    if (-not (Get-Command vidmuse -ErrorAction SilentlyContinue)) {
        $tmp = Join-Path $env:TEMP "install-vidmuse.ps1"
        Invoke-WebRequest "https://vidmuse.sandcdn.com/cli/install.ps1" -OutFile $tmp -UseBasicParsing
        $installedPath = (& powershell -ExecutionPolicy Bypass -File $tmp | Select-Object -Last 1)
        $installedDir = Split-Path -Parent $installedPath
        if (($env:PATH -split [System.IO.Path]::PathSeparator) -notcontains $installedDir) {
            $env:PATH = "$installedDir$([System.IO.Path]::PathSeparator)$env:PATH"
        }
    }
    vidmuse --version
    
  2. Authenticate when needed.

    First check whether the current session is already authenticated:

    vidmuse profile get --output json
    

    If the command returns an auth error, ask before starting any login flow because browser and device login require user action.

    Use this decision tree:

    Environment Command
    Local desktop with browser access vidmuse login
    Headless or remote terminal with streaming output vidmuse login --device
    Agent shell cannot keep a long-running command visible vidmuse login --device --start, show the URL/code to the user, then run vidmuse login --device --complete after the user confirms authorization
  3. For command routing, read references/command-map.md.

Local media runtime

The released CLI binary does not require Go. Account, model, asset, style, voice, thread, message, and memory commands require only the vidmuse binary and network access. Check extra tools before using local preview or rendering commands:

node --version
ffmpeg -version
Command Required runtime
vidmuse serve FFmpeg only when video thumbnails are requested
vidmuse render Node.js 22+ and FFmpeg

When Node.js or FFmpeg is missing, do not install system packages without the user's approval. Tell the user which prerequisite is missing and offer the appropriate command:

# macOS (Homebrew)
brew install node ffmpeg

# Debian or Ubuntu
curl -fsSL https://deb.nodesource.com/setup_22.x | sudo -E bash -
sudo apt-get install -y nodejs ffmpeg

The render command uses npx --yes hyperframes@0.7.26 when no HYPERFRAMES_BIN is supplied. The first such render needs npm network access. Use NODE_BIN, FFMPEG_BIN, and HYPERFRAMES_BIN for nonstandard executable paths.

Output Contract

Default output mode is text. Add --output json only when structured parsing is needed. Do not add it to vidmuse render: its -o/--output flag specifies the rendered video file path.

Use text for compact plain-language responses. Nested lists and maps are rendered as indented blocks.

Use table only for short human-facing summaries because top-level structs, maps, lists, and lists of maps are rendered as tables when possible, while nested lists and maps stay compact inside table cells.

Paginated text/table output uses a human summary such as:

Showing 1-20 of 88

JSON output preserves structured fields such as limit, offset, total, and data when the API returns them.

Streams:

Stream Meaning
stdout Command result in the selected output mode
stderr Errors, verbose HTTP logs, installation logs

Exit codes:

Code Meaning
0 Success
1 General error
2 Usage or validation error
3 Auth error: missing, expired, or invalid auth
4 Network or API error

Error handling:

  1. Check the exit code first.
  2. If exit code is 0, parse stdout only when json mode was requested.
  3. If exit code is non-zero, show a concise error based on stderr.
  4. For exit code 3, ask the user to run vidmuse login or vidmuse login --device.
  5. For exit code 4, mention the network or API failure and retry only when the operation is read-only or explicitly safe to repeat.

Safety

Read-only commands may be run directly when the user asks for information. Mutating commands such as thread create, message send, memory create, memory update, memory append, memory push, and memory pop require explicit user intent.

Never print auth tokens, cookies, or full config files in the final answer because final answers can be saved to transcripts, shared as screenshots, or copied into bug reports.

Related skills

video-editgenmedia-labs715KEdit existing video on RunComfy — this skill is a smart router that matches the user's intent to the right edit model in the RunComfy catalog. Picks Wan 2.7 Edit-Video (general restyle / background swap / packaging swap, identity + motion preservation), Kling 2.6 Pro Motion Control (transfer precise motion from a reference video to a target character), or Lucy Edit Restyle (lightweight identity-stable restyle / outfit swap). Bundles each model's documented prompting patterns so the skill gets shai-video-generationgenmedia-labs714KGenerate AI videos on RunComfy via the `runcomfy` CLI — a smart router across the full video-model catalog: HappyHorse 1.0 (Arena #1, native in-pass audio), Wan-AI Wan 2-7 (open weights, audio-driven lip-sync), ByteDance Seedance v2 / 1-5 / 1-0 (multi-modal cinematic), Kling 3.0 / 2-6, Google Veo 3-1, MiniMax Hailuo 2-3, ByteDance Dreamina 3-0. Covers text-to-video (t2v), image-to-video (i2v), and Veo's video-extend endpoint. The skill picks the right model for the user's intent (Arena-#1 qualitai-musicgenmedia-labs714KGenerate AI music on RunComfy via the `runcomfy` CLI — a smart router across the music-model catalog. Routes to ElevenLabs AI Music Generation (premium 44.1 kHz stereo vocal tracks, 5 s–5 min, $0.0083/s) and ACE Step / ACE Step 1.5 (StepFun-AI open-weights, tag-driven composition, multilingual lyrics, $0.0002–0.0003/s, ~27× cheaper), plus ACE Step audio-inpaint (regenerate a time range inside an existing track) and ACE Step audio-outpaint (extend a track before or after). Picks the right model fimage-to-videogenmedia-labs713KAnimate any still image on RunComfy — this skill is a smart router that matches the user's intent to the right i2v model in the RunComfy catalog. Picks HappyHorse 1.0 I2V (Arena #1, native audio, identity preservation) for general animations, Wan 2.7 with `audio_url` for custom-voiceover lip-sync, or Seedance 2.0 Pro for multi-modal animation from image + reference video + reference audio. Bundles each model's documented prompting patterns so the caller gets sharper output without burning iterat

Search skills and MCP servers

Fuzzy search across 23,137 skills and servers