Agent Skills

agnes-ai-generation

videoyacey2.3K installs

Call Agnes AI / Sapiens AI generation APIs for text, image, and video. Use when the user asks to use Agnes models, Agnes Image, Agnes Video, Agnes 2.0 Flash, apihub.agnes-ai.com, or to generate text, images, edit images, create videos, animate images, create keyframe videos, or test Agnes API calls.

Install

npx skills add https://github.com/yacey/agnes-ai-generation-skill --skill agnes-ai-generation
SKILL.md

Agnes AI Generation

Use this skill to call Agnes text, image, and video generation APIs through https://apihub.agnes-ai.com.

Quick Start

  1. Read references/api.md when endpoint details, parameters, or response fields are needed.
  2. Use scripts/agnes_api.py for real API calls instead of rewriting curl by hand.
  3. Require an API key in AGNES_API_KEY, AGNES_API_TOKEN, or APIHUB_AGNES_API_KEY. Never print the key.
  4. For light live verification, run smoke-test; it avoids video creation by default. Add --include-image-edit for image-to-image, and add --video-case <case> explicitly for video modes. Treat the skill as fully tested only when basic text, text streaming, text tool calling, text-to-image, image-to-image, text-to-video, image-to-video, multi-image video, keyframe video, and video retrieval return successful responses.

Commands

Text generation:

python scripts/agnes_api.py text --prompt "Write a concise product tagline for an AI assistant."

Streaming text:

python scripts/agnes_api.py text --prompt "Write a short product intro." --stream

Streaming output is normalized and includes aggregated content, events, done, and a short raw_prefix.

Image generation:

python scripts/agnes_api.py image --prompt "A luminous floating city above a misty canyon at sunrise, cinematic realism" --size 1024x768

Image-to-image:

python scripts/agnes_api.py image --prompt "Turn the scene into a rainy cyberpunk night while preserving composition" --image https://example.com/input.png --size 1024x768

Text-to-video with polling:

python scripts/agnes_api.py video --prompt "A cinematic shot of a cat walking on the beach at sunset" --poll

Image-to-video:

python scripts/agnes_api.py video --prompt "Animate subtle camera movement and natural lighting" --image https://example.com/image.png --poll

Keyframe / multi-image video:

python scripts/agnes_api.py video --prompt "Create a smooth cinematic transition between the two keyframes" --image https://example.com/a.png --image https://example.com/b.png --mode keyframes --poll

Retrieve a video task:

python scripts/agnes_api.py video-get video_123456

Light live smoke test:

python scripts/agnes_api.py smoke-test

Image edit smoke test:

python scripts/agnes_api.py smoke-test --include-image-edit

Single video smoke test:

python scripts/agnes_api.py smoke-test --video-case text-to-video

Workflow

  • Prefer agnes-2.0-flash for text chat/completions.
  • Do not use Agnes Responses API multi-turn function calling for autonomous tool workflows. Live testing showed the provider can return function_call with overall status=completed, and submitting function_call_output with previous_response_id may fail. Use this skill's chat completions path for text generation and treat tool-calling as best-effort request-shape compatibility only.
  • Prefer agnes-image-2.1-flash for text-to-image, image-to-image, and high-information-density image generation. High-density generation is prompt-driven; include subject hierarchy, environment, secondary details, lighting, composition, and quality requirements.
  • Prefer agnes-video-v2.0 for text-to-video, image-to-video, multi-image video, keyframe animation, prompt-based motion and scene control, cinematic output, asynchronous task creation, polling-based result retrieval, and seed-based reproducibility.
  • For image and video generation, convert any non-English user prompt to a fluent English generation prompt before calling the image/video API. English prompts are more stable for Agnes video generation. Preserve concrete visual details, style, lighting, composition, motion, camera instructions, and constraints during translation.
  • For videos, remember the API is asynchronous: create a task first, then poll or retrieve by video_id when the create response includes it. The script falls back to legacy task_id lookup only when video_id is absent.
  • The script validates image sizes, video frame counts, frame rates, and dimensions before sending requests. num_frames must be 8n + 1 and <= 441; 81 or 121 are good short values.
  • The video command defaults to num_frames=121 and frame_rate=24 for more stable generation. Video smoke tests default to num_frames=81 and frame_rate=24.
  • Warn the user before costly or long-running live video generation unless they explicitly asked to test or generate video.
  • Test video capabilities one at a time with smoke-test --video-case <case> to avoid creating many tasks at once. Supported cases are text-to-video, image-to-video, multi-image, and keyframes.

Current Validation Notes

  • Confirmed locally: skill metadata validation and Python syntax.
  • Confirmed by live API: basic text, streaming text, tool-calling request shape, text-to-image, image-to-image, high-information-density text-to-image, Chinese prompt translation for image/video, completed text-to-video URL retrieval, and completed image-to-video URL retrieval.
  • Caveat: Agnes may accept tool-calling request parameters without consistently returning tool_calls; use smoke-test --strict-tools when strict tool-call validation is required.
  • Caveat: Agnes Responses API multi-turn function calling is not reliable for agent tool loops; do not rely on it for Codex/Claude-style automatic tool continuation.
  • Supported by the script and smoke-test selector, but not re-run end-to-end in the latest pass: multi-image video and keyframe animation.
  • Not yet confirmed end-to-end: completed URL retrieval for every multi-image video and keyframe animation task. A previous text-to-video task returned a provider-side division by zero error, so keep video retries visible and report provider errors clearly.

Output Handling

  • Return generated image/video URLs directly by default. Do not download, save, open, or inspect generated media unless the user explicitly asks for a local file or visual inspection.
  • For image responses, expect URL-style results when extra_body.response_format is url.
  • For video responses, extract URLs from video_url, url, or remixed_from_video_id when status is completed.
  • For video retrieval, prefer GET /agnesapi?video_id=...&model_name=agnes-video-v2.0; legacy GET /v1/videos/{task_id} remains a fallback.
  • If a request fails, report HTTP status and provider error body without exposing the API key.

Related skills

video-editgenmedia-labs715KEdit existing video on RunComfy — this skill is a smart router that matches the user's intent to the right edit model in the RunComfy catalog. Picks Wan 2.7 Edit-Video (general restyle / background swap / packaging swap, identity + motion preservation), Kling 2.6 Pro Motion Control (transfer precise motion from a reference video to a target character), or Lucy Edit Restyle (lightweight identity-stable restyle / outfit swap). Bundles each model's documented prompting patterns so the skill gets shai-video-generationgenmedia-labs714KGenerate AI videos on RunComfy via the `runcomfy` CLI — a smart router across the full video-model catalog: HappyHorse 1.0 (Arena #1, native in-pass audio), Wan-AI Wan 2-7 (open weights, audio-driven lip-sync), ByteDance Seedance v2 / 1-5 / 1-0 (multi-modal cinematic), Kling 3.0 / 2-6, Google Veo 3-1, MiniMax Hailuo 2-3, ByteDance Dreamina 3-0. Covers text-to-video (t2v), image-to-video (i2v), and Veo's video-extend endpoint. The skill picks the right model for the user's intent (Arena-#1 qualitai-musicgenmedia-labs714KGenerate AI music on RunComfy via the `runcomfy` CLI — a smart router across the music-model catalog. Routes to ElevenLabs AI Music Generation (premium 44.1 kHz stereo vocal tracks, 5 s–5 min, $0.0083/s) and ACE Step / ACE Step 1.5 (StepFun-AI open-weights, tag-driven composition, multilingual lyrics, $0.0002–0.0003/s, ~27× cheaper), plus ACE Step audio-inpaint (regenerate a time range inside an existing track) and ACE Step audio-outpaint (extend a track before or after). Picks the right model fimage-to-videogenmedia-labs713KAnimate any still image on RunComfy — this skill is a smart router that matches the user's intent to the right i2v model in the RunComfy catalog. Picks HappyHorse 1.0 I2V (Arena #1, native audio, identity preservation) for general animations, Wan 2.7 with `audio_url` for custom-voiceover lip-sync, or Seedance 2.0 Pro for multi-modal animation from image + reference video + reference audio. Bundles each model's documented prompting patterns so the caller gets sharper output without burning iterat

Search skills and MCP servers

Fuzzy search across 23,137 skills and servers