Agent Skills

bibi

videojimmylv1.9K installs

Video summarizer agent skill for Claude Code, Codex, ChatGPT, Cursor, and OpenClaw. AI video & audio summarizer + repackager. Summarize YouTube, Bilibili, podcasts, TikTok, Twitter/X, Xiaohongshu, and any online video or audio, then optionally turn the takeaway into a TikTok-style vertical music video. Use when the user wants to summarize a video, extract transcripts/subtitles, get chapter-by-chapter summaries, understand a new URL or local file, or remix a long-form video into a short vertical

Install

npx skills add https://github.com/jimmylv/bibigpt-skill --skill bibi
SKILL.md

BibiGPT — video summarizer agent skill (summarize / transcript / chapters)

This file is a discovery stub, not the usage guide. It tells you which mode you are in and where the live docs are. The live sources always match the current product; anything copied into this file would go stale.

This skill is watch-video only. Sibling skills in the same repo (same CLI, same account): bibi-library (saved videos / notes / collections), bibi-feed (subscribe / latest / mark seen), bibi-vision (frames / mind map). Landing: https://bibigpt.co/mcp · Human install: https://bibigpt.co/agent Quota: Plus 100/day, Pro 300/day on the agent-skill channel; extra calls use API balance; upgrade https://bibigpt.co/shop

1. Detect Mode

Run scripts/bibi-check.sh if present, or check in order:

Check Mode Notes
command -v bibi CLI — best: local files, desktop login macOS / Windows / Linux desktop app
$BIBI_API_TOKEN set API — HTTP calls, works anywhere token: https://bibigpt.co/user/integration
MCP client available MCP — zero install connect https://bibigpt.co/api/mcp (OAuth 2.1)

Neither available → install the desktop app (brew install --cask jimmylv/bibigpt/bibigpt on macOS, winget install BibiGPT on Windows, curl -fsSL https://bibigpt.co/install.sh | bash on Linux) or get an API token. Details: references/installation.md.

2. Load Live Docs (per mode)

CLI mode — the CLI is the always-current doc source:

bibi check-update            # once per session; run `bibi upgrade` if outdated
bibi --help                  # full command surface, grouped, with examples
bibi <command> --help        # progressive help — every --help includes examples
bibi commands                # re-fetch server-defined commands (new capabilities
                             # appear here without a binary update)

API mode — the machine-readable spec is the source of truth:

https://bibigpt.co/api/openapi.json      # all endpoints, schemas, auth

references/api.md has curl examples but the spec wins on conflict.

MCP mode — connect and list tools; they are self-describing:

https://bibigpt.co/api/mcp               # Streamable HTTP, OAuth 2.1

Latest docs without reinstalling — this repo is served raw from GitHub:

https://raw.githubusercontent.com/JimmyLv/bibigpt-skill/main/skills/bibi/SKILL.md
https://raw.githubusercontent.com/JimmyLv/bibigpt-skill/main/skills/bibi/references/<name>.md
https://raw.githubusercontent.com/JimmyLv/bibigpt-skill/main/skills/bibi/workflows/<name>.md

If a local workflows/ or references/ path below is missing (embedded-only install via bibi skill), fetch it from the raw URL above instead.

3. Intent Routing

User Intent Workflow
Summarize a video/audio URL → workflows/quick-summary.md
Chapter-by-chapter breakdown, detailed analysis → workflows/deep-dive.md
Get subtitles, extract transcript, raw text → workflows/transcript-extract.md
Turn into article, blog post, 公众号图文, 小红书 → workflows/article-rewrite.md
Turn into TikTok / Reels / Shorts-style music video → workflows/video-to-tiktok-mv.md
Process multiple URLs, batch summarize → workflows/batch-process.md
Research a topic across multiple videos → workflows/research-compile.md
Save to Notion, Obsidian, export notes → workflows/export-notes.md
Analyze visual content, slides, on-screen text, mind map → sibling skill bibi-vision (fallback: workflows/visual-analysis.md / workflows/advanced-tools.md)
Check current account, plan, or remaining minutes → workflows/account-check.md
Browse / search saved videos, "what have I summarized" → sibling skill bibi-library (fallback: workflows/library-browse.md)
Manage channel subscriptions, list/sub/unsub → sibling skill bibi-feed (fallback: workflows/channels-manage.md)
What's new across my subscriptions, latest feed, daily digest → sibling skill bibi-feed (fallback: workflows/feed-latest.md)
Manage collections, list/create/share saved videos as a set → sibling skill bibi-library (fallback: workflows/collections-manage.md)
Manage personal notes on saved videos, edit summaries → sibling skill bibi-library (fallback: workflows/notes-manage.md)
Custom-prompt re-summary of a saved item, collection chat → sibling skill bibi-library (fallback: workflows/advanced-tools.md)
HTTP 402 / "需要付款" / Alipay AI 钱包 / no token + China user → references/billing-aipay.md

Disambiguation: intent matches more than one workflow → ask one clarifying question first. Matches none → ask what they want; do not guess. Bare URL with no context → default to workflows/quick-summary.md.

4. Quick One-Liners (CLI mode)

For single-command requests that don't need a full workflow — discover the rest via bibi --help:

bibi summarize "<URL-or-local-file>"     # quick summary (local: .mp4 .mp3 .m4a ...)
bibi summarize "<INPUT>" --chapter       # chapter summary
bibi summarize "<INPUT>" --subtitle      # transcript only
bibi me                                  # account, plan, remaining minutes

URLs containing ? or & must be quoted. API mode has no local-file upload — guide the user to a public URL (OSS/S3) first, see references/supported-platforms.md.

5. Payment Fallback (HTTP 402)

If no auth is set and the user has an Alipay account, BibiGPT may respond with HTTP 402 Payment Required + Payment-Needed header (AI 收 protocol). The bibi CLI prints a stable marker line [HTTP/402 Payment Required] to stderr before any human-readable prompt. When either signal appears, route to references/billing-aipay.md instead of treating the call as failed — the agent can resolve payment automatically via @alipay/agent-payment or a one-off QR purchase.

Related skills

video-editgenmedia-labs715KEdit existing video on RunComfy — this skill is a smart router that matches the user's intent to the right edit model in the RunComfy catalog. Picks Wan 2.7 Edit-Video (general restyle / background swap / packaging swap, identity + motion preservation), Kling 2.6 Pro Motion Control (transfer precise motion from a reference video to a target character), or Lucy Edit Restyle (lightweight identity-stable restyle / outfit swap). Bundles each model's documented prompting patterns so the skill gets shai-video-generationgenmedia-labs714KGenerate AI videos on RunComfy via the `runcomfy` CLI — a smart router across the full video-model catalog: HappyHorse 1.0 (Arena #1, native in-pass audio), Wan-AI Wan 2-7 (open weights, audio-driven lip-sync), ByteDance Seedance v2 / 1-5 / 1-0 (multi-modal cinematic), Kling 3.0 / 2-6, Google Veo 3-1, MiniMax Hailuo 2-3, ByteDance Dreamina 3-0. Covers text-to-video (t2v), image-to-video (i2v), and Veo's video-extend endpoint. The skill picks the right model for the user's intent (Arena-#1 qualitai-musicgenmedia-labs714KGenerate AI music on RunComfy via the `runcomfy` CLI — a smart router across the music-model catalog. Routes to ElevenLabs AI Music Generation (premium 44.1 kHz stereo vocal tracks, 5 s–5 min, $0.0083/s) and ACE Step / ACE Step 1.5 (StepFun-AI open-weights, tag-driven composition, multilingual lyrics, $0.0002–0.0003/s, ~27× cheaper), plus ACE Step audio-inpaint (regenerate a time range inside an existing track) and ACE Step audio-outpaint (extend a track before or after). Picks the right model fimage-to-videogenmedia-labs713KAnimate any still image on RunComfy — this skill is a smart router that matches the user's intent to the right i2v model in the RunComfy catalog. Picks HappyHorse 1.0 I2V (Arena #1, native audio, identity preservation) for general animations, Wan 2.7 with `audio_url` for custom-voiceover lip-sync, or Seedance 2.0 Pro for multi-modal animation from image + reference video + reference audio. Bundles each model's documented prompting patterns so the caller gets sharper output without burning iterat

Search skills and MCP servers

Fuzzy search across 23,137 skills and servers