Agent Skills

heygen

videoheygen-com1.7K installs

Install

npx skills add https://github.com/heygen-com/skills --skill heygen
SKILL.md

HeyGen Video Agent — NanoClaw Container Skill

When to Use

Use this skill when the user wants to create a video with an AI avatar presenter. Triggers: "make a video", "create a video message", "record a video", "avatar video", "talking head video", "video pitch", "video update".

NOT for: image generation, audio-only TTS, video translation, or cinematic b-roll.

Required Environment

Steps

Step 1: Discover Available Avatars

heygen avatar list --ownership public --limit 5 | jq '.data[] | {group_id: .id, avatar_name: .name}'

avatar list returns avatar groups — each .id is a group_id, not an avatar_id. The avatar_id you pass to generation is a specific look: list looks with heygen avatar looks list --group-id <group_id> | jq '.data[] | {avatar_id: .id, preview_image_url}' and pick a look's .id. If the user already has a specific look id, use it directly.

Step 2: Find a Voice

heygen voice list --limit 10 | jq '.data[] | {voice_id, name, language}'

Pick a voice_id matching the desired language and tone.

Step 3: Write the Script

Write a spoken-word script for the avatar. Rules:

  • Write for speech, not text. Short sentences. Natural pauses.
  • 150 words per minute is the target pace.
  • 30 seconds = ~75 words. 60 seconds = ~150 words.
  • No stage directions. No markdown. Just what the avatar says.

Step 4: Generate the Video

heygen video-agent create \
  --prompt "YOUR SCRIPT HERE" \
  --avatar-id "AVATAR_ID_HERE" \
  --voice-id "VOICE_ID_HERE" \
  --orientation landscape \
  --wait --timeout 10m

With --wait, the CLI blocks until the video completes and emits the final status object. Without --wait, submission returns immediately — save the session_id from stdout for manual polling.

--timeout 10m is sized for short videos (≤60s). Bump to --timeout 45m for videos >60s or when targeting 2+ minute output — Video Agent generation routinely takes 20-45m for longer clips.

Step 5: Poll for Completion (only without --wait)

heygen video-agent get SESSION_ID | jq '{status: .data.status, video_id: .data.video_id}'

Poll every 15 seconds. Session status is one of thinking, waiting_for_input, reviewing, generating, completed, failed — not a strict linear sequence (waiting_for_input only occurs in chat mode). Terminal states are completed and failed.

Once .data.video_id is present, run heygen video get <video_id> and read .data.video_url for the download URL (and .data.failure_message on failure).

Step 6: Deliver

Download the video and present it to the user:

heygen video download <video_id>

Writes the MP4 to disk and emits {"asset", "message", "path"} on stdout — chain on .path.

Verification

After generating a video, confirm:

  1. CLI exits 0 (generation accepted). Without --wait, stdout includes a session_id for polling; with --wait, the CLI polls the video to completion and stdout is the final video resource (.data.id, .data.video_url).
  2. Polling (or --wait) returns status: "completed" within 5 minutes
  3. heygen video get <video_id> returns a valid HTTPS .data.video_url
  4. Downloaded file is a playable MP4

Report Feedback

After the video is delivered and verified, report how the run went (once per task):

heygen feedback --rating <1-5> --comment "<what worked, or what broke>"

--rating runs 1 (broke / unusable) to 5 (worked great); put any bug or friction in --comment. Anonymous, no API key; no-ops if analytics are opted out (HEYGEN_NO_ANALYTICS or heygen config set analytics false).

Troubleshooting

Symptom Fix
Exit code 3 / auth error on stderr Check heygen auth status; run heygen auth login or set HEYGEN_API_KEY
Exit code 2 / usage error Run heygen video-agent create --help — verify flag names and required args
Status stuck on thinking / generating Wait up to 5 minutes. Videos over 60s take longer.
Missing video_id Session may have failed. Check .data.status; if failed, inspect the full heygen video-agent get <session_id> response for the failure detail.

Limits

  • Free tier: 1 minute of video per month
  • API trial: 3 free credits on signup
  • Max video length per request: ~5 minutes
  • Concurrent generation: depends on plan tier

Related skills

video-editgenmedia-labs715KEdit existing video on RunComfy — this skill is a smart router that matches the user's intent to the right edit model in the RunComfy catalog. Picks Wan 2.7 Edit-Video (general restyle / background swap / packaging swap, identity + motion preservation), Kling 2.6 Pro Motion Control (transfer precise motion from a reference video to a target character), or Lucy Edit Restyle (lightweight identity-stable restyle / outfit swap). Bundles each model's documented prompting patterns so the skill gets shai-video-generationgenmedia-labs714KGenerate AI videos on RunComfy via the `runcomfy` CLI — a smart router across the full video-model catalog: HappyHorse 1.0 (Arena #1, native in-pass audio), Wan-AI Wan 2-7 (open weights, audio-driven lip-sync), ByteDance Seedance v2 / 1-5 / 1-0 (multi-modal cinematic), Kling 3.0 / 2-6, Google Veo 3-1, MiniMax Hailuo 2-3, ByteDance Dreamina 3-0. Covers text-to-video (t2v), image-to-video (i2v), and Veo's video-extend endpoint. The skill picks the right model for the user's intent (Arena-#1 qualitai-musicgenmedia-labs714KGenerate AI music on RunComfy via the `runcomfy` CLI — a smart router across the music-model catalog. Routes to ElevenLabs AI Music Generation (premium 44.1 kHz stereo vocal tracks, 5 s–5 min, $0.0083/s) and ACE Step / ACE Step 1.5 (StepFun-AI open-weights, tag-driven composition, multilingual lyrics, $0.0002–0.0003/s, ~27× cheaper), plus ACE Step audio-inpaint (regenerate a time range inside an existing track) and ACE Step audio-outpaint (extend a track before or after). Picks the right model fimage-to-videogenmedia-labs713KAnimate any still image on RunComfy — this skill is a smart router that matches the user's intent to the right i2v model in the RunComfy catalog. Picks HappyHorse 1.0 I2V (Arena #1, native audio, identity preservation) for general animations, Wan 2.7 with `audio_url` for custom-voiceover lip-sync, or Seedance 2.0 Pro for multi-modal animation from image + reference video + reference audio. Bundles each model's documented prompting patterns so the caller gets sharper output without burning iterat

Search skills and MCP servers

Fuzzy search across 23,137 skills and servers