Agent Skills

Generate videos with Google Veo models via inference.sh CLI. Models: Veo 3.1, Veo 3.1 Fast. Capabilities: text-to-video, cinematic output, high quality video generation. Triggers: veo, google veo, veo 3, veo 2, veo 3.1, vertex ai video, google video generation, google video ai, veo model, veo video

Install

npx skills add https://github.com/inference-sh/skills --skill google-veo
SKILL.md

Install the belt CLI skill: npx skills add belt-sh/cli

Google Veo Video Generation

Generate videos with Google Veo models via inference.sh CLI.

Google Veo Video Generation

Quick Start

Requires inference.sh CLI (belt). Install instructions

belt login

belt app run google/veo-3-1-fast --input '{"prompt": "drone shot over a mountain lake"}'

Veo Models

Model App ID Speed Quality
Veo 3.1 google/veo-3-1 Slower Best
Veo 3.1 Fast google/veo-3-1-fast Fast Excellent

Search Veo Apps

belt app search "veo"

Examples

Cinematic Shot

belt app run google/veo-3-1-fast --input '{
  "prompt": "Cinematic drone shot flying through a misty forest at sunrise, volumetric lighting"
}'

Product Demo

belt app run google/veo-3-1 --input '{
  "prompt": "Sleek smartphone rotating on a dark reflective surface, studio lighting"
}'

Nature Scene

belt app run google/veo-3-1-fast --input '{
  "prompt": "Timelapse of clouds moving over a mountain range, golden hour"
}'

Action Shot

belt app run google/veo-3-1 --input '{
  "prompt": "Slow motion water droplet splashing into a pool, macro shot"
}'

Urban Scene

belt app run google/veo-3-1-fast --input '{
  "prompt": "Busy city street at night with neon signs and rain reflections, Tokyo style"
}'

Prompt Tips

Camera movements: drone shot, tracking shot, pan, zoom, dolly, steadicam

Lighting: golden hour, blue hour, studio lighting, volumetric, neon, natural

Style: cinematic, documentary, commercial, artistic, realistic

Timing: slow motion, timelapse, real-time

Sample Workflow

# 1. Generate sample input to see all options
belt app sample google/veo-3-1-fast --save input.json

# 2. Edit the prompt
# 3. Run
belt app run google/veo-3-1-fast --input input.json

Related Skills

# Full platform skill (all apps)
npx skills add inference-sh/skills@infsh-cli

# All video generation models
npx skills add inference-sh/skills@ai-video-generation

# AI avatars & lipsync
npx skills add inference-sh/skills@ai-avatar-video

# Image generation (for image-to-video)
npx skills add inference-sh/skills@ai-image-generation

Browse all video apps: belt app list --category video

Documentation

Related skills

video-editgenmedia-labs715KEdit existing video on RunComfy — this skill is a smart router that matches the user's intent to the right edit model in the RunComfy catalog. Picks Wan 2.7 Edit-Video (general restyle / background swap / packaging swap, identity + motion preservation), Kling 2.6 Pro Motion Control (transfer precise motion from a reference video to a target character), or Lucy Edit Restyle (lightweight identity-stable restyle / outfit swap). Bundles each model's documented prompting patterns so the skill gets shai-video-generationgenmedia-labs714KGenerate AI videos on RunComfy via the `runcomfy` CLI — a smart router across the full video-model catalog: HappyHorse 1.0 (Arena #1, native in-pass audio), Wan-AI Wan 2-7 (open weights, audio-driven lip-sync), ByteDance Seedance v2 / 1-5 / 1-0 (multi-modal cinematic), Kling 3.0 / 2-6, Google Veo 3-1, MiniMax Hailuo 2-3, ByteDance Dreamina 3-0. Covers text-to-video (t2v), image-to-video (i2v), and Veo's video-extend endpoint. The skill picks the right model for the user's intent (Arena-#1 qualitai-musicgenmedia-labs714KGenerate AI music on RunComfy via the `runcomfy` CLI — a smart router across the music-model catalog. Routes to ElevenLabs AI Music Generation (premium 44.1 kHz stereo vocal tracks, 5 s–5 min, $0.0083/s) and ACE Step / ACE Step 1.5 (StepFun-AI open-weights, tag-driven composition, multilingual lyrics, $0.0002–0.0003/s, ~27× cheaper), plus ACE Step audio-inpaint (regenerate a time range inside an existing track) and ACE Step audio-outpaint (extend a track before or after). Picks the right model fimage-to-videogenmedia-labs713KAnimate any still image on RunComfy — this skill is a smart router that matches the user's intent to the right i2v model in the RunComfy catalog. Picks HappyHorse 1.0 I2V (Arena #1, native audio, identity preservation) for general animations, Wan 2.7 with `audio_url` for custom-voiceover lip-sync, or Seedance 2.0 Pro for multi-modal animation from image + reference video + reference audio. Bundles each model's documented prompting patterns so the caller gets sharper output without burning iterat

Search skills and MCP servers

Fuzzy search across 23,137 skills and servers