Agent Skills

image-generation

Create effective AI image generation prompts for DALL-E, Midjourney, and Stable Diffusion. Generate prompts for various styles and use cases.

Install

npx skills add https://github.com/claude-office-skills/skills --skill image-generation
SKILL.md

Image Generation Skill

Overview

I help you create effective prompts for AI image generation tools like DALL-E, Midjourney, and Stable Diffusion. I understand the nuances of different platforms and can help you achieve specific visual styles.

What I can do:

  • Write detailed image generation prompts
  • Optimize prompts for specific AI tools
  • Suggest style keywords and modifiers
  • Create negative prompts to avoid unwanted elements
  • Adapt prompts for different aspect ratios
  • Generate variations and alternatives

What I cannot do:

  • Generate images directly
  • Guarantee exact output from AI tools
  • Predict how AI will interpret prompts
  • Bypass content policies of AI tools

How to Use Me

Step 1: Describe Your Vision

Tell me:

  • What you want to see in the image
  • The purpose (presentation, social media, marketing)
  • Style preferences (realistic, artistic, minimalist)
  • Mood or emotion to convey

Step 2: Choose the Platform

  • DALL-E 3: Best for clarity and instruction-following
  • Midjourney: Best for artistic and aesthetic images
  • Stable Diffusion: Most customizable, local options

Step 3: Specify Parameters

  • Aspect ratio (1:1, 16:9, 4:3, etc.)
  • Quality level
  • Style references
  • Things to avoid

Prompt Engineering Framework

Basic Prompt Structure

[Subject] + [Action/State] + [Environment] + [Style] + [Technical Parameters]

Detailed Template

[Main Subject]
- Who/what is the focus?
- What are they doing?

[Environment/Setting]
- Where is this taking place?
- Time of day? Weather? Season?

[Composition]
- Camera angle (eye-level, bird's eye, low angle)
- Framing (close-up, medium shot, wide shot)
- Focus (depth of field)

[Style]
- Art style (photorealistic, watercolor, oil painting, etc.)
- Artist reference (optional)
- Era/period

[Lighting]
- Type (natural, studio, dramatic, soft)
- Direction (backlit, side-lit, front-lit)

[Color]
- Palette (warm, cool, monochrome)
- Specific colors to include

[Mood/Atmosphere]
- Emotion to evoke
- Overall feeling

[Technical]
- Quality modifiers
- Aspect ratio
- Negative prompts

Platform-Specific Tips

DALL-E 3

Strengths:

  • Follows complex instructions well
  • Good at text in images
  • Natural language understanding

Prompt Style: Write naturally, be descriptive

A professional photograph of a modern office space with floor-to-ceiling 
windows overlooking a city skyline at sunset. The room features a minimalist 
wooden desk with a laptop, a potted monstera plant, and warm ambient lighting 
from a designer floor lamp. The mood is productive yet peaceful.

Midjourney

Strengths:

  • Exceptional aesthetics
  • Strong artistic styles
  • Good at specific looks

Prompt Style: Use keywords, parameters, and style references

modern office space, floor-to-ceiling windows, city skyline sunset, 
minimalist wooden desk, monstera plant, warm ambient lighting, 
productive atmosphere --ar 16:9 --style raw --v 6

Key Parameters:

  • --ar X:Y Aspect ratio
  • --v 6 Version
  • --style raw Less stylized
  • --q 2 Quality
  • --s 250 Stylize amount

Stable Diffusion

Strengths:

  • Highly customizable
  • Local/private generation
  • Extensive community models

Prompt Style: Weighted tokens, negative prompts essential

(modern office:1.2), (floor-to-ceiling windows:1.1), city skyline, 
golden hour sunset, minimalist wooden desk, laptop, monstera plant, 
(warm ambient lighting:1.3), professional photograph, 8k, detailed

Negative: cartoon, drawing, illustration, (worst quality:1.4), 
(low quality:1.4), blurry, watermark

Style Keywords Reference

Photography Styles

Style Keywords
Portrait portrait photography, headshot, bokeh, 85mm lens
Product product photography, studio lighting, white background
Landscape landscape photography, golden hour, dramatic sky
Street street photography, candid, urban, documentary
Fashion fashion editorial, vogue style, high fashion

Art Styles

Style Keywords
Realistic photorealistic, hyperrealistic, 8k, detailed
Illustration digital illustration, vector art, flat design
Watercolor watercolor painting, soft edges, flowing colors
Oil Painting oil painting, brush strokes, impasto
Anime anime style, manga, cel shading
3D Render 3D render, octane render, blender, CGI

Mood/Atmosphere

Mood Keywords
Professional corporate, business, clean, modern
Cozy warm, inviting, comfortable, hygge
Dramatic cinematic, high contrast, moody lighting
Cheerful bright, colorful, happy, vibrant
Minimalist simple, clean, whitespace, zen

Lighting

Type Keywords
Natural natural light, soft daylight, golden hour
Studio studio lighting, softbox, rim light
Dramatic chiaroscuro, dramatic shadows, volumetric
Neon neon lights, cyberpunk, colorful glow

Output Format

# Image Generation Prompts: [Concept]

**Purpose**: [What the image is for]
**Target Platform**: [DALL-E / Midjourney / SD]
**Aspect Ratio**: [X:Y]

---

## Primary Prompt

### For DALL-E 3:

[Natural language prompt with full description]


### For Midjourney:

[Keyword-based prompt with parameters]


### For Stable Diffusion:

[Weighted prompt]

Negative: [Things to avoid]


---

## Variations

### Variation 1: [Description]

[Prompt]


### Variation 2: [Description]

[Prompt]


### Variation 3: [Description]

[Prompt]


---

## Style Alternatives

### Option A: [Style Name]
Add these keywords: `[keywords]`

### Option B: [Style Name]
Add these keywords: `[keywords]`

---

## Tips for This Image

1. [Specific tip for achieving desired result]
2. [Tip]
3. [Tip]

---

## Iteration Suggestions

If the result isn't quite right:
- Try: [Adjustment 1]
- Try: [Adjustment 2]
- Try: [Adjustment 3]

Common Use Cases

Business/Corporate

Professional headshot, corporate portrait, business casual attire, 
neutral background, studio lighting, confident expression, 
sharp focus, high resolution

Marketing/Social Media

Lifestyle product photography, natural setting, soft natural light, 
pastel color palette, Instagram aesthetic, flat lay composition, 
millennial pink, minimalist

Presentation Graphics

Abstract business concept, isometric illustration, blue and white 
color scheme, clean design, corporate style, flat design, 
technology theme, professional

Blog/Article Headers

Wide banner image, [topic] concept art, editorial style, 
cinematic composition, 16:9 aspect ratio, headline-friendly 
(space for text overlay), muted colors

Tips for Better Results

  1. Be specific - "golden retriever" > "dog"
  2. Describe what you want, not what you don't (use negative prompts separately)
  3. Use quality modifiers - "professional", "detailed", "8k"
  4. Reference styles - "in the style of [artist/genre]"
  5. Specify composition - "close-up", "wide shot", "bird's eye view"
  6. Include lighting - dramatically affects mood
  7. Iterate - refine based on results
  8. Use consistent seeds for variations (when possible)

Limitations

  • Cannot generate images directly
  • Results vary between AI platforms
  • Some concepts are restricted by platform policies
  • Exact output is unpredictable
  • May require multiple iterations

Built by the Claude Office Skills community. Contributions welcome!

Related skills

ai-image-generationgenmedia-labs713KGenerate and edit images on RunComfy via the `runcomfy` CLI — a smart router across the full image-model catalog: FLUX 2 (Klein 9B/4B, Pro, Dev, Flash, Turbo, Max), Google Nano Banana 2 / Pro, OpenAI GPT Image 2, ByteDance Seedream 5 / 4-5 / 4-0 and Dreamina 4-0, Alibaba Qwen Image and Z-Image Turbo, Wan 2-7. Covers both text-to-image (t2i) and image-to-image / edit (i2i) endpoints — the skill picks the right model for the user's actual intent (typography precision, photoreal portraits, sub-secoai-image-generation101-skills547KGenerate AI images with GPT-Image-2, FLUX, Gemini, Grok, Seedream, Reve and 50+ models via inference.sh CLI. Models: GPT-Image-2, FLUX Dev LoRA, FLUX.2 Klein LoRA, Gemini 3 Pro Image, Grok Imagine, Seedream 4.5, Reve, ImagineArt. Capabilities: text-to-image, image-to-image, inpainting, LoRA, image editing, upscaling, text rendering. Use for: AI art, product mockups, concept art, social media graphics, marketing visuals, illustrations. Triggers: flux, image generation, ai image, text to image, stnano-banana-2prime-skills424KGenerate images with Google Nano Banana 2 (Gemini-family flash-tier text-to-image) on RunComfy — bundled with the model's documented prompting patterns so the skill gets sharper output than naive prompting against the same model. Documents Nano Banana 2's strengths (rapid iteration, in-image typography rendering, predictable framing, optional web-grounded context), the resolution-tier pricing, the safety-tolerance dial, and when to route to Nano Banana Pro / GPT Image 2 / Flux 2 / Seedream insteimage-editprime-skills424KEdit images on RunComfy — this skill is a smart router that matches the user's intent to the right edit model in the RunComfy catalog. Picks Nano Banana Edit (batch up to 20, identity-preserving default), OpenAI GPT Image 2 Edit (multilingual in-image text rewrite, multi-ref composition, layout precision), Flux Kontext Pro (single-ref high-fidelity local edit), or Z-Image Turbo Inpaint (mask-driven precise region edit). Bundles each model's documented prompting patterns so the skill gets sharper

Search skills and MCP servers

Fuzzy search across 23,137 skills and servers