xAI Grok image and video generation guide covering authentication, endpoints, prompt structure, image editing, reference-image video, and async polling.
Install
npx skills add https://github.com/calesthio/openmontage --skill grok-mediaSKILL.md
Grok Media
Use this skill when working with xAI media models in OpenMontage.
Models
grok-imagine-imagefor image generation and image editinggrok-imagine-videofor text-to-video, image-to-video, and reference-image video
Authentication
- Env var:
XAI_API_KEY - Base URL:
https://api.x.ai/v1 - Header:
Authorization: Bearer $XAI_API_KEY
Image API
Text-to-image
- Endpoint:
POST /images/generations - Core fields:
modelpromptnaspect_ratioresolution
Image edit
- Endpoint:
POST /images/edits - Use
imagefor one source image - Use
imagesfor multi-image compositing - Each source image can be:
- a public HTTPS URL
- a base64 data URI
Image prompting
- Grok responds well to direct natural language
- For edits, describe only the intended change and preserve everything else implicitly
- For multi-image merges, explicitly name how each source contributes
- Prefer one strong scene description over long style-stacking
Video API
Generation
- Endpoint:
POST /videos/generations - Polling endpoint:
GET /videos/{request_id} - Success state:
status == "done" - Failure states to handle explicitly:
failed,expired
Modes
- Text-to-video:
- prompt-only generation
- Image-to-video:
- use
image: {"url": ...} - this anchors the starting frame
- use
- Reference-to-video:
- use
reference_images: [{"url": ...}, ...] - this influences who/what appears in the video without locking the first frame
- prompts can reference inputs with placeholders like
<IMAGE_1>,<IMAGE_2>
- use
Video constraints
- Grok video is best treated as short-form generation
- Current output resolutions are
480pand720p - Reference-image video supports multiple images and is useful for product placement, wardrobe transfer, and identity consistency
- Download outputs promptly; provider URLs may be temporary
Pricing
grok-imagine-image:$0.02per generated imagegrok-imagine-imageedits/composites: add$0.002per input imagegrok-imagine-video:480p:$0.05per second720p:$0.07per second
grok-imagine-videoimage-conditioned requests: add$0.002per input image
Grok-Specific Prompt Guidance
Images
- Start with subject, action, setting
- Add one style anchor, not five
- For edits:
- describe the desired modification
- keep the rest of the image stable by omission, not by writing a giant preservation list
Video
- Keep prompts scene-local: one shot, one main motion idea, one emotional beat
- For reference-conditioned video, explicitly map source images to roles:
- person from
<IMAGE_1> - jacket from
<IMAGE_2> - product from
<IMAGE_3>
- person from
- Camera and pacing language helps:
- slow push-in
- handheld follow
- locked-off medium shot
- high-energy whip pan transition
Good Fits
- Image style transfer
- Image compositing from multiple sources
- Reference-conditioned short video
- Product-led motion clips
- Character-consistent scenes without hard first-frame lock
Weak Fits
- Long-form clip generation
- Heavy reliance on deterministic seeds
- Overloaded prompts with multiple scene changes
Failure Handling
- If generation submission succeeds but polling expires, surface it as a provider/runtime issue
- If a request fails, preserve the endpoint, mode, and prompt summary in the error
- Do not silently substitute a different provider after xAI was selected without user approval
Related skills
ai-image-generationgenmedia-labs713KGenerate and edit images on RunComfy via the `runcomfy` CLI — a smart router across the full image-model catalog: FLUX 2 (Klein 9B/4B, Pro, Dev, Flash, Turbo, Max), Google Nano Banana 2 / Pro, OpenAI GPT Image 2, ByteDance Seedream 5 / 4-5 / 4-0 and Dreamina 4-0, Alibaba Qwen Image and Z-Image Turbo, Wan 2-7. Covers both text-to-image (t2i) and image-to-image / edit (i2i) endpoints — the skill picks the right model for the user's actual intent (typography precision, photoreal portraits, sub-secoai-image-generation101-skills547KGenerate AI images with GPT-Image-2, FLUX, Gemini, Grok, Seedream, Reve and 50+ models via inference.sh CLI. Models: GPT-Image-2, FLUX Dev LoRA, FLUX.2 Klein LoRA, Gemini 3 Pro Image, Grok Imagine, Seedream 4.5, Reve, ImagineArt. Capabilities: text-to-image, image-to-image, inpainting, LoRA, image editing, upscaling, text rendering. Use for: AI art, product mockups, concept art, social media graphics, marketing visuals, illustrations. Triggers: flux, image generation, ai image, text to image, stnano-banana-2prime-skills424KGenerate images with Google Nano Banana 2 (Gemini-family flash-tier text-to-image) on RunComfy — bundled with the model's documented prompting patterns so the skill gets sharper output than naive prompting against the same model. Documents Nano Banana 2's strengths (rapid iteration, in-image typography rendering, predictable framing, optional web-grounded context), the resolution-tier pricing, the safety-tolerance dial, and when to route to Nano Banana Pro / GPT Image 2 / Flux 2 / Seedream insteimage-editprime-skills424KEdit images on RunComfy — this skill is a smart router that matches the user's intent to the right edit model in the RunComfy catalog. Picks Nano Banana Edit (batch up to 20, identity-preserving default), OpenAI GPT Image 2 Edit (multilingual in-image text rewrite, multi-ref composition, layout precision), Flux Kontext Pro (single-ref high-fidelity local edit), or Z-Image Turbo Inpaint (mask-driven precise region edit). Bundles each model's documented prompting patterns so the skill gets sharper
