从图片生成白板手绘动画视频。将任意彩色图片转换为包含线稿绘制和上色两个阶段的动画,带手部覆盖效果,输出 H.264 MP4 视频。支持单张和批量两种模式。当用户说"把图片做成白板动画"、"白板动画"、"批量白板动画"时触发。
Install
npx skills add https://github.com/yangagent/whiteboard-animation-skill --skill whiteboard-animationSKILL.md
白板手绘动画生成器
从输入图片生成白板手绘动画视频,动画分为两个阶段:
- 线稿绘制 — 手持笔在白板上逐步画出黑白线稿
- 上色 — 手持笔沿内容轮廓逐步涂上彩色,还原为原图
支持两种模式:单张模式和批量模式。
模式判断
- 用户提供单张图片路径 → 使用单张模式
- 用户提供多张图片路径(数组/列表) → 使用批量模式
单张模式工作流
第一步:准备环境
先用 --check 检测环境是否就绪:
python <skill目录>/scripts/setup_env.py --check
- 如果成功(退出码 0):最后一行输出
PYTHON_PATH=<路径>,捕获该路径用于后续步骤,直接跳到第二步 - 如果失败(退出码 1):运行完整安装:
python <skill目录>/scripts/setup_env.py
安装脚本会自动创建 .venv 虚拟环境并安装缺失依赖(opencv-python、numpy、av),最后一行同样输出 PYTHON_PATH=<路径>。
第二步:确认输入图片
从用户请求中获取图片路径,确认文件存在。支持格式:PNG、JPG、JPEG、BMP、TIFF。
白色或浅色背景的图片效果最佳。
第三步:确定参数
收集可选参数,所有参数都有合理的默认值:
| 参数 | 标志 | 默认值 | 说明 |
|---|---|---|---|
| 图片路径 | 位置参数(必填) | -- | 输入的彩色图片路径 |
| 输出目录 | --output-dir |
./output |
视频输出目录 |
| 时长 | --duration |
10000 |
视频总时长(毫秒) |
| 无手部 | --no-hand |
默认显示手 | 禁用手部覆盖效果 |
第四步:运行生成脚本
使用第一步获取的 PYTHON_PATH 运行生成脚本:
<PYTHON_PATH> <skill目录>/scripts/generate_whiteboard.py <图片路径> [--output-dir <目录>] [--duration <毫秒>] [--no-hand]
示例:
<PYTHON_PATH> <skill目录>/scripts/generate_whiteboard.py /path/to/photo.png --output-dir ./output --duration 20000
第五步:返回结果
脚本会将最终视频路径打印到 stdout,将该路径告知用户。输出文件命名格式:vid_YYYYMMDD_HHMMSS_h264.mp4。
批量模式工作流
当用户提供多张图片(图片路径数组)和对应的时长数组时,使用批量模式。
第一步:准备环境
与单张模式相同。先运行 setup_env.py 获取 PYTHON_PATH。
第二步:校验输入
从用户请求中获取:
- 图片路径数组:多张图片的路径列表(必填)
- 时长数组:与图片一一对应的时长列表(毫秒,必填)
必须满足:
- 图片路径数组和时长数组长度相同
- 每张图片文件都存在
- 每个时长为正整数(毫秒)
第三步:运行批量生成脚本
使用 PYTHON_PATH 运行批量脚本,通过 --images 和 --durations 传入一一对应的图片路径和时长:
<PYTHON_PATH> <skill目录>/scripts/batch_generate.py \
--images /path/to/img1.png /path/to/img2.png /path/to/img3.png \
--durations 10000 15000 8000 \
[--output-dir ./output] [--no-hand]
脚本内部会自动校验:
--images和--durations数量必须一致- 每张图片文件必须存在
校验通过后串行逐个调用单张生成脚本,每完成一个打印进度。
第四步:返回结果
脚本输出所有任务的汇总信息(成功/失败数量)。告知用户:
- 总共生成了多少个视频
- 成功几个,失败几个
- 输出目录路径
- 如有失败,列出失败的图片
故障排除
ModuleNotFoundError:重新运行setup_env.py确保依赖完整安装。- 虚拟环境创建失败:确认系统已安装 Python 3.8+,且
python3命令可用。 - 批量模式单个任务失败:不影响后续任务继续执行,最终汇总会列出所有失败项。
Related skills
video-editgenmedia-labs715KEdit existing video on RunComfy — this skill is a smart router that matches the user's intent to the right edit model in the RunComfy catalog. Picks Wan 2.7 Edit-Video (general restyle / background swap / packaging swap, identity + motion preservation), Kling 2.6 Pro Motion Control (transfer precise motion from a reference video to a target character), or Lucy Edit Restyle (lightweight identity-stable restyle / outfit swap). Bundles each model's documented prompting patterns so the skill gets shai-video-generationgenmedia-labs714KGenerate AI videos on RunComfy via the `runcomfy` CLI — a smart router across the full video-model catalog: HappyHorse 1.0 (Arena #1, native in-pass audio), Wan-AI Wan 2-7 (open weights, audio-driven lip-sync), ByteDance Seedance v2 / 1-5 / 1-0 (multi-modal cinematic), Kling 3.0 / 2-6, Google Veo 3-1, MiniMax Hailuo 2-3, ByteDance Dreamina 3-0. Covers text-to-video (t2v), image-to-video (i2v), and Veo's video-extend endpoint. The skill picks the right model for the user's intent (Arena-#1 qualitai-musicgenmedia-labs714KGenerate AI music on RunComfy via the `runcomfy` CLI — a smart router across the music-model catalog. Routes to ElevenLabs AI Music Generation (premium 44.1 kHz stereo vocal tracks, 5 s–5 min, $0.0083/s) and ACE Step / ACE Step 1.5 (StepFun-AI open-weights, tag-driven composition, multilingual lyrics, $0.0002–0.0003/s, ~27× cheaper), plus ACE Step audio-inpaint (regenerate a time range inside an existing track) and ACE Step audio-outpaint (extend a track before or after). Picks the right model fimage-to-videogenmedia-labs713KAnimate any still image on RunComfy — this skill is a smart router that matches the user's intent to the right i2v model in the RunComfy catalog. Picks HappyHorse 1.0 I2V (Arena #1, native audio, identity preservation) for general animations, Wan 2.7 with `audio_url` for custom-voiceover lip-sync, or Seedance 2.0 Pro for multi-modal animation from image + reference video + reference audio. Bundles each model's documented prompting patterns so the caller gets sharper output without burning iterat