Agent Skills

Apply AI-powered analysis to images for business-specific tagging, metadata extraction, and quality checks using controlled vocabularies. Use when user wants to analyze images and apply structured metadata in ImageKit.

Install

npx skills add https://github.com/imagekit-developer/skills --skill ai-tasks
SKILL.md

AI Tasks Skill

When to use

  • User wants to tag images with business-specific categories
  • User needs controlled vocabularies (predefined value lists)
  • User wants yes/no quality checks on images
  • User needs to extract metadata using natural language analysis

AI Task Structure

Each AI Task contains 1-10 sub-tasks with:

  1. Instruction (required): Natural language question about the image
  2. Action Type (required): select_tags, select_metadata, or yes_no
  3. Vocabulary (optional): 1-30 predefined approved values

Action Types

select_tags

Adds tags from vocabulary. Supports multiple selections.

{
  "type": "select_tags",
  "instruction": "What body style is this vehicle?",
  "vocabulary": ["sedan", "suv", "hatchback"],
  "max_selections": 1
}

select_metadata

Sets custom metadata field (field must exist in DAM).

{
  "type": "select_metadata",
  "instruction": "What is the primary color?",
  "field": "primary_color",
  "vocabulary": ["red", "blue", "white", "black"],
  "max_selections": 1
}

yes_no

Binary quality check with conditional actions.

{
  "type": "yes_no",
  "instruction": "Is the product completely visible?",
  "on_yes": { "add_tags": ["framing_ok"] },
  "on_no": { "add_tags": ["needs_reshot"] }
}

Gotchas

  • Instruction clarity: Be specific ("What is the collar type?" not "Describe the image")
  • Vocabulary design: Use business terminology, non-overlapping, 1-30 items max
  • Field requirements: For select_metadata, field must exist in DAM first
  • Tag values: Cannot contain % character
  • yes_no tasks: Must have at least one of on_yes or on_no
  • Vocabulary length: Max 500 characters combined (select_tags only)
  • Scope: Works on visual content only (images/videos)

Applying AI Tasks

  • Via DAM MCP: list_saved_extensions, create_saved_extension, and apply_extension_bulk on https://imagekit.io/mcp/dam
  • Via Saved Extensions: Create and apply via dashboard/API
  • Via API at Upload: Include in extensions array
  • Via Path Policies: Auto-apply to files in specific folders

Full examples

For complete, copy-ready ai-tasks configurations organized by industry (fashion e-commerce, travel, automotive) and detailed per-task-type parameter references, read resources/EXAMPLES.md.

Related skills

ai-image-generationgenmedia-labs713KGenerate and edit images on RunComfy via the `runcomfy` CLI — a smart router across the full image-model catalog: FLUX 2 (Klein 9B/4B, Pro, Dev, Flash, Turbo, Max), Google Nano Banana 2 / Pro, OpenAI GPT Image 2, ByteDance Seedream 5 / 4-5 / 4-0 and Dreamina 4-0, Alibaba Qwen Image and Z-Image Turbo, Wan 2-7. Covers both text-to-image (t2i) and image-to-image / edit (i2i) endpoints — the skill picks the right model for the user's actual intent (typography precision, photoreal portraits, sub-secoai-image-generation101-skills547KGenerate AI images with GPT-Image-2, FLUX, Gemini, Grok, Seedream, Reve and 50+ models via inference.sh CLI. Models: GPT-Image-2, FLUX Dev LoRA, FLUX.2 Klein LoRA, Gemini 3 Pro Image, Grok Imagine, Seedream 4.5, Reve, ImagineArt. Capabilities: text-to-image, image-to-image, inpainting, LoRA, image editing, upscaling, text rendering. Use for: AI art, product mockups, concept art, social media graphics, marketing visuals, illustrations. Triggers: flux, image generation, ai image, text to image, stnano-banana-2prime-skills424KGenerate images with Google Nano Banana 2 (Gemini-family flash-tier text-to-image) on RunComfy — bundled with the model's documented prompting patterns so the skill gets sharper output than naive prompting against the same model. Documents Nano Banana 2's strengths (rapid iteration, in-image typography rendering, predictable framing, optional web-grounded context), the resolution-tier pricing, the safety-tolerance dial, and when to route to Nano Banana Pro / GPT Image 2 / Flux 2 / Seedream insteimage-editprime-skills424KEdit images on RunComfy — this skill is a smart router that matches the user's intent to the right edit model in the RunComfy catalog. Picks Nano Banana Edit (batch up to 20, identity-preserving default), OpenAI GPT Image 2 Edit (multilingual in-image text rewrite, multi-ref composition, layout precision), Flux Kontext Pro (single-ref high-fidelity local edit), or Z-Image Turbo Inpaint (mask-driven precise region edit). Bundles each model's documented prompting patterns so the skill gets sharper

Search skills and MCP servers

Fuzzy search across 23,137 skills and servers