Agent Skills

music

videoelevenlabs7.6K installs

Generate music using ElevenLabs Music API. Use when creating instrumental tracks, songs with lyrics, background music, jingles, or any AI-generated music composition. Supports prompt-based generation, composition plans for granular control, and detailed output with metadata.

Install

npx skills add https://github.com/elevenlabs/skills --skill music
SKILL.md

ElevenLabs Music Generation

Generate music from text prompts - supports instrumental tracks, songs with lyrics, and fine-grained control via composition plans.

Setup: See Installation Guide. For JavaScript, use @elevenlabs/* packages only.

All examples below use music_v2_5, the most advanced generation model. Pass music_v2 or music_v1 only when an older model is explicitly requested.

Quick Start

Python

from elevenlabs import ElevenLabs

client = ElevenLabs()

audio = client.music.compose(
    prompt="A chill lo-fi hip hop beat with jazzy piano chords",
    music_length_ms=30000,
    model_id="music_v2_5",
)

with open("output.mp3", "wb") as f:
    for chunk in audio:
        f.write(chunk)

TypeScript

import { ElevenLabsClient } from "@elevenlabs/elevenlabs-js";
import { createWriteStream } from "fs";

const client = new ElevenLabsClient();
const audio = await client.music.compose({
  prompt: "A chill lo-fi hip hop beat with jazzy piano chords",
  musicLengthMs: 30000,
  modelId: "music_v2_5",
});
audio.pipe(createWriteStream("output.mp3"));

CLI

elevenlabs music compose \
  --prompt "A chill lo-fi beat" \
  --music-length-ms 30000 \
  --model-id music_v2_5 \
  --output output.mp3

Methods

Method Description
music.compose Generate audio from a prompt or composition plan
music.stream Stream audio chunks as they are generated (paid plans)
music.composition_plan.create Generate a structured plan for fine-grained control
music.compose_detailed Generate audio + composition plan + metadata; pass store_for_inpainting=True to enable inpainting
music.compose_detailed_stream Stream audio plus composition plan, metadata, and optional word timestamps as Server-Sent Events
music.video_to_music Generate background music from one or more uploaded video files
music.upload Upload an audio file for later inpainting workflows, optionally extracting its composition plan or word-level timestamps
music.finetunes.list List accessible music finetunes
music.finetunes.create Train a music finetune from uploaded audio
music.finetunes.get Retrieve finetune status and metadata
music.finetunes.update Update finetune metadata or visibility
music.finetunes.delete Delete a music finetune

See API Reference for full parameter details.

music.upload is available to enterprise clients with access to the inpainting feature.

Music Finetunes

Create a finetune from training audio with POST /v1/music/finetunes, then poll the get endpoint until its status is completed. Pass the returned id as finetune_id when composing music.

Use the list, update, and delete endpoints to manage accessible finetunes.

Video to Music

Generate background music from uploaded video clips via POST /v1/music/video-to-music (client.music.video_to_music). This is separate from prompt-based music.compose (POST /v1/music).

The API combines videos in order, accepts an optional natural-language description, and lets you steer style with up to 10 tags such as upbeat or cinematic. This endpoint still defaults to music_v1; pass model_id="music_v2_5" to use the most advanced model.

Python

from elevenlabs import ElevenLabs

client = ElevenLabs()

audio = client.music.video_to_music(
    videos=["trailer.mp4"],
    description="Build suspense, then resolve with a warm cinematic finish.",
    tags=["cinematic", "suspenseful", "uplifting"],
    model_id="music_v2_5",
)

with open("video-score.mp3", "wb") as f:
    for chunk in audio:
        f.write(chunk)

TypeScript

import { ElevenLabsClient } from "@elevenlabs/elevenlabs-js";
import { createReadStream, createWriteStream } from "fs";

const client = new ElevenLabsClient();

const audio = await client.music.videoToMusic({
  videos: [createReadStream("trailer.mp4")],
  description: "Build suspense, then resolve with a warm cinematic finish.",
  tags: ["cinematic", "suspenseful", "uplifting"],
  modelId: "music_v2_5",
});

audio.pipe(createWriteStream("video-score.mp3"));

CLI

elevenlabs music video_to_music \
  --videos trailer.mp4 \
  --description "Build suspense, then resolve with a warm cinematic finish." \
  --tags cinematic \
  --model-id music_v2_5 \
  --output video-score.mp3

The CLI currently accepts one --videos file and one --tags value per request; use the Python or TypeScript SDK to send multiple videos or tags.

Constraints from the current API schema:

  • Upload 1-10 video files per request
  • Keep total combined upload size at or below 200 MB
  • Keep total combined video duration at or below 600 seconds
  • Use description for high-level musical direction and tags for concise style cues

Composition Plans

music_v2_5 composition plans are an ordered list of chunks. Each chunk specifies its own text (section label, lyrics, inline cues), duration_ms, positive_styles, negative_styles, and context_adherence (low, medium, or high, default high). Up to 30 chunks per plan, each 3,000–120,000 ms, total length 3 s to 10 minutes. Each chunk's text supports up to 6,132 characters, with up to 30 lines of 200 characters each.

Generate a plan first, edit it, then compose:

plan = client.music.composition_plan.create(
    prompt="An epic orchestral piece building to a climax",
    music_length_ms=60000,
    model_id="music_v2_5",
)

# Edit chunks in place
plan["chunks"][0]["text"] = "[Intro]\nQuiet strings rising"

audio = client.music.compose(
    composition_plan=plan,
    model_id="music_v2_5",
)
const plan = await client.music.compositionPlan.create({
  prompt: "An epic orchestral piece building to a climax",
  musicLengthMs: 60000,
  modelId: "music_v2_5",
});

plan.chunks[0].text = "[Intro]\nQuiet strings rising";

const audio = await client.music.compose({
  compositionPlan: plan,
  modelId: "music_v2_5",
});

Or hand-build a plan to control lyrics and style per section:

composition_plan = {
    "chunks": [
        {
            "text": "[Verse]\nWalking down an empty street",
            "duration_ms": 15000,
            "positive_styles": ["pop", "upbeat", "female vocals", "acoustic guitar"],
            "negative_styles": ["dark", "slow"],
            "context_adherence": "high",
        },
        {
            "text": "[Chorus]\nThis is my moment",
            "duration_ms": 15000,
            "positive_styles": ["powerful vocals", "full band"],
            "negative_styles": [],
            "context_adherence": "high",
        },
    ]
}

audio = client.music.compose(composition_plan=composition_plan, model_id="music_v2_5")
const compositionPlan = {
  chunks: [
    {
      text: "[Verse]\nWalking down an empty street",
      durationMs: 15000,
      positiveStyles: ["pop", "upbeat", "female vocals", "acoustic guitar"],
      negativeStyles: ["dark", "slow"],
      contextAdherence: "high",
    },
    {
      text: "[Chorus]\nThis is my moment",
      durationMs: 15000,
      positiveStyles: ["powerful vocals", "full band"],
      negativeStyles: [],
      contextAdherence: "high",
    },
  ],
};

const audio = await client.music.compose({
  compositionPlan,
  modelId: "music_v2_5",
});

Put broader characteristics (genre, instrumentation, vocal style) in positive_styles, not in text. The first chunk's styles set the overall tone — include 6–7 styles there.

Output Formats

Use the output_format query parameter on compose, detailed compose, or stream requests to select the generated audio format. auto chooses a model-appropriate MP3 format; for music_v2 and music_v2_5, it selects mp3_48000_192. Higher-bitrate MP3 options include mp3_48000_240 and mp3_48000_320.

Streaming

For paid plans, stream audio chunks as they are generated instead of waiting for the full file:

from io import BytesIO

stream = client.music.stream(
    prompt="A driving synthwave track with arpeggiated leads",
    music_length_ms=30000,
    model_id="music_v2_5",
)

buffer = BytesIO()
for chunk in stream:
    if chunk:
        buffer.write(chunk)
const stream = await client.music.stream({
  prompt: "A driving synthwave track with arpeggiated leads",
  musicLengthMs: 30000,
  modelId: "music_v2_5",
});

const chunks: Buffer[] = [];
for await (const chunk of stream) {
  chunks.push(chunk);
}

Detailed streaming

Use detailed streaming when the application needs generated music metadata while audio is still arriving. POST /v1/music/detailed/stream accepts the same prompt or composition-plan body as detailed compose, streams text/event-stream, and can include word timestamps with with_timestamps.

elevenlabs music compose_detailed_stream \
  --prompt "A bright indie pop hook with warm guitars" \
  --music-length-ms 30000 \
  --model-id music_v2_5 \
  --with-timestamps true \
  --output-format auto

Inpainting

Inpainting edits or extends a stored song by mixing audio reference chunks (unchanged slices of a stored song) with new generation chunks in a single composition plan.

Step 1 — get a song_id, either by storing a fresh generation or uploading existing audio:

# Option A: keep a generation for later editing
result = client.music.compose_detailed(
    prompt="An upbeat pop song with verse and chorus",
    music_length_ms=60000,
    model_id="music_v2_5",
    store_for_inpainting=True,
)
song_id = result.song_id

# Option B: upload an existing track and extract its plan
uploaded = client.music.upload(
    file=open("my-song.mp3", "rb"),
    extract_composition_plan="music_v2_5",
)
song_id = uploaded.song_id
composition_plan = uploaded.composition_plan
import { createReadStream } from "fs";

// Option A: keep a generation for later editing
const result = await client.music.composeDetailed({
  prompt: "An upbeat pop song with verse and chorus",
  musicLengthMs: 60000,
  modelId: "music_v2_5",
  storeForInpainting: true,
});
let songId = result.songId;

// Option B: upload an existing track and extract its plan
const uploaded = await client.music.upload({
  file: createReadStream("my-song.mp3"),
  extractCompositionPlan: "music_v2_5",
});
songId = uploaded.songId;
const compositionPlan = uploaded.compositionPlan;

Step 2 — compose a plan that references the stored audio and regenerates the part you want to change:

plan = {
    "chunks": [
        {"song_id": song_id, "range": {"start_ms": 0, "end_ms": 30000}},
        {
            "text": "[Chorus]\nWe're rising up tonight",
            "duration_ms": 30000,
            "positive_styles": ["bigger drums", "layered vocals", "anthemic"],
            "negative_styles": ["sparse"],
            "context_adherence": "high",
        },
    ]
}

audio = client.music.compose(composition_plan=plan, model_id="music_v2_5")
const plan = {
  chunks: [
    { songId, range: { startMs: 0, endMs: 30000 } },
    {
      text: "[Chorus]\nWe're rising up tonight",
      durationMs: 30000,
      positiveStyles: ["bigger drums", "layered vocals", "anthemic"],
      negativeStyles: ["sparse"],
      contextAdherence: "high",
    },
  ],
};

const audio = await client.music.compose({
  compositionPlan: plan,
  modelId: "music_v2_5",
});

To match the feel of a stored slice without copying it, attach a conditioning_ref (up to 30,000 ms) plus a condition_strength of low, medium, high, or xhigh to a generation chunk. Conditioning placed on the first chunk influences every later chunk.

See API Reference for the full inpainting parameter list.

Content Restrictions

  • Cannot reference specific artists, bands, or copyrighted lyrics
  • bad_prompt errors include a prompt_suggestion with alternative phrasing
  • bad_composition_plan errors include a composition_plan_suggestion

Error Handling

try:
    audio = client.music.compose(prompt="...", music_length_ms=30000)
except Exception as e:
    print(f"API error: {e}")
try {
  const audio = await client.music.compose({
    prompt: "...",
    musicLengthMs: 30000,
  });
} catch (err) {
  console.error("API error:", err);
}

Common errors: 401 (invalid key), 422 (invalid params), 429 (rate limit).

References

Related skills

video-editgenmedia-labs715KEdit existing video on RunComfy — this skill is a smart router that matches the user's intent to the right edit model in the RunComfy catalog. Picks Wan 2.7 Edit-Video (general restyle / background swap / packaging swap, identity + motion preservation), Kling 2.6 Pro Motion Control (transfer precise motion from a reference video to a target character), or Lucy Edit Restyle (lightweight identity-stable restyle / outfit swap). Bundles each model's documented prompting patterns so the skill gets shai-video-generationgenmedia-labs714KGenerate AI videos on RunComfy via the `runcomfy` CLI — a smart router across the full video-model catalog: HappyHorse 1.0 (Arena #1, native in-pass audio), Wan-AI Wan 2-7 (open weights, audio-driven lip-sync), ByteDance Seedance v2 / 1-5 / 1-0 (multi-modal cinematic), Kling 3.0 / 2-6, Google Veo 3-1, MiniMax Hailuo 2-3, ByteDance Dreamina 3-0. Covers text-to-video (t2v), image-to-video (i2v), and Veo's video-extend endpoint. The skill picks the right model for the user's intent (Arena-#1 qualitai-musicgenmedia-labs714KGenerate AI music on RunComfy via the `runcomfy` CLI — a smart router across the music-model catalog. Routes to ElevenLabs AI Music Generation (premium 44.1 kHz stereo vocal tracks, 5 s–5 min, $0.0083/s) and ACE Step / ACE Step 1.5 (StepFun-AI open-weights, tag-driven composition, multilingual lyrics, $0.0002–0.0003/s, ~27× cheaper), plus ACE Step audio-inpaint (regenerate a time range inside an existing track) and ACE Step audio-outpaint (extend a track before or after). Picks the right model fimage-to-videogenmedia-labs713KAnimate any still image on RunComfy — this skill is a smart router that matches the user's intent to the right i2v model in the RunComfy catalog. Picks HappyHorse 1.0 I2V (Arena #1, native audio, identity preservation) for general animations, Wan 2.7 with `audio_url` for custom-voiceover lip-sync, or Seedance 2.0 Pro for multi-modal animation from image + reference video + reference audio. Bundles each model's documented prompting patterns so the caller gets sharper output without burning iterat

Search skills and MCP servers

Fuzzy search across 23,137 skills and servers