---
title: "Music generation"
canonical: https://docs.qvac.tether.io/sdk/ai-capabilities/music-generation/
collection: "SDK"
package: "@qvac/sdk"
line: v0.21
current_line: true
---

# Music generation (/sdk/ai-capabilities/music-generation)



## Overview

Music generation uses [`@qvac/audiogen-ggml`](https://github.com/tetherto/qvac/tree/main/packages/audiogen-ggml), with your choice of either [ACE-Step 1.5](https://github.com/ace-step/ACE-Step-1.5) or [MiniMax-Music3](https://huggingface.co/MiniMaxAI/MiniMax-Music3) as the inference engine. Load a model with `modelType: "audiogen"` and select the engine with `modelConfig.engine` — `"minimax"` for MiniMax-Music3; ACE-Step is the default. Then call `audioGen()` with a caption and any engine-specific inputs.

Pick ACE-Step when you need lyrics, per-song musical controls (BPM, key, time signature, vocal language), reference-audio timbre conditioning, cover generation, [editing of an existing recording](#edit-an-existing-recording) with `audioEdit()`, or [describing one](#understand-a-recording) with `audioUnderstand()`. Pick MiniMax-Music3 when you want a lighter API surface — just a caption and a handful of sampling knobs — on desktop hardware with a capable GPU.

`audioGen()` (and `audioEdit()`) returns immediately with:

* `requestId` — a synchronous identifier for targeted cancellation.
* `progressStream` — stage and step updates while generation runs.
* `audio` — a promise resolving to interleaved PCM, its sample rate, channel count, and `bitsPerSample`.
* `stats` — a promise resolving to timing statistics.

Both engines return through this same shape; only the inputs to `loadModel()` and `audioGen()` differ. `audioUnderstand()` runs the same pipeline backwards and so returns `description` where the other two return `audio`.

<Callout type="info">
  **Platform support:** ACE-Step runs on desktop (Linux, macOS, Windows) and
  mobile (Android arm64, iOS arm64), with GPU acceleration via Metal (Apple) and
  Vulkan (Android). MiniMax-Music3 runs on desktop only and needs \~22 GB of
  device memory for its f16 model pair when `useGPU: true`.
</Callout>

## Functions

Use the following sequence:

1. [`loadModel()`](/sdk/reference/api#loadmodel)
2. [`audioGen()`](/sdk/reference/api#audiogen) — or [`audioEdit()`](/sdk/reference/api#audioedit) to edit an existing recording, or [`audioUnderstand()`](/sdk/reference/api#audiounderstand) to describe one (both ACE-Step only)
3. [`unloadModel()`](/sdk/reference/api#unloadmodel)

For complete signatures, see the [SDK API reference](/sdk/reference/api/).

## Enable the plugin

When your project selects plugins explicitly, add the AudioGen plugin:

```json title="qvac.config.json"
{
  "plugins": ["@qvac/sdk/audiogen-ggml/plugin"]
}
```

Then rebuild the SDK bundle:

```bash
qvac bundle sdk
```

If `plugins` is omitted or empty, all built-in plugins are included. See the [plugin system](/sdk/configuration/plugins) for bundle and Bare runtime registration details.

## Models

Unlike single-file model families, AudioGen does not use a top-level `modelSrc`. Supply every model source in `modelConfig`, and select the engine with `modelConfig.engine`.

### ACE-Step

ACE-Step uses four GGUF files — text encoder, LM, one DiT variant of the three listed, and VAE:

| Stage          | SDK model constant                   | Purpose                              |
| -------------- | ------------------------------------ | ------------------------------------ |
| Text encoder   | `AUDIOGEN_QWEN3_EMBEDDING_0_6B_Q8_0` | Encodes captions and lyrics          |
| Language model | `AUDIOGEN_ACESTEP_5HZ_LM_0_6B_Q8_0`  | Plans the song and musical structure |
| DiT            | `AUDIOGEN_ACESTEP_V15_TURBO_Q4_K_M`  | Fast, lower-memory Turbo generation  |
| DiT            | `AUDIOGEN_ACESTEP_V15_TURBO_Q8_0`    | Higher-precision Turbo generation    |
| DiT            | `AUDIOGEN_ACESTEP_V15_SFT_Q8_0`      | Supervised fine-tuned generation     |
| VAE            | `AUDIOGEN_VAE_BF16`                  | Decodes the latent into PCM audio    |

<Callout type="warning">
  The first ACE-Step load downloads \~3.3 GB across these four files. On slow or
  unstable links, set `registryStreamTimeoutMs: 600000` and
  `registryDownloadMaxRetries: 10` in `qvac.config.json` and point
  `QVAC_CONFIG_PATH` at that file before loading.
</Callout>

```ts
import {
  AUDIOGEN_ACESTEP_5HZ_LM_0_6B_Q8_0,
  AUDIOGEN_ACESTEP_V15_TURBO_Q4_K_M,
  AUDIOGEN_QWEN3_EMBEDDING_0_6B_Q8_0,
  AUDIOGEN_VAE_BF16,
  loadModel,
} from "@qvac/sdk";

const modelId = await loadModel({
  modelType: "audiogen",
  modelConfig: {
    engine: "acestep",
    textEncModelSrc: AUDIOGEN_QWEN3_EMBEDDING_0_6B_Q8_0,
    lmModelSrc: AUDIOGEN_ACESTEP_5HZ_LM_0_6B_Q8_0,
    ditModelSrc: AUDIOGEN_ACESTEP_V15_TURBO_Q4_K_M,
    vaeModelSrc: AUDIOGEN_VAE_BF16,
    useGPU: true,
  },
});
```

Advanced load-time controls (`inferenceSteps`, `shift`, `nGpuLayers`, `threads`, `backendsDir`) are ACE-Step's to tune; omit `inferenceSteps` and `shift` unless you need to override the addon's per-variant defaults. `backendsDir` is needed on arm64, where the CPU backend ships as per-microarch modules.

For more about model constants and sources, see [SDK — Models](/sdk#models).

### MiniMax-Music3

MiniMax-Music3 uses two GGUF files that you supply as local paths — the language model (`mm3-lm-*.gguf`) and the synthesis model (`mm3-synth-*.gguf`). The SDK does not download MiniMax-Music3 from the registry, so bring your own converted files and point `lmModelSrc` and `synthModelSrc` at them. Weights are governed by the [MiniMax-Music3 Community License](https://huggingface.co/MiniMaxAI/MiniMax-Music3).

```ts
import { loadModel } from "@qvac/sdk";

const modelId = await loadModel({
  modelType: "audiogen",
  modelConfig: {
    engine: "minimax",
    lmModelSrc: "/path/to/mm3-lm-q8.gguf",
    synthModelSrc: "/path/to/mm3-synth-q8.gguf",
    useGPU: true,
  },
});
```

## Generate audio

`caption` is required and cannot be empty. Everything else is engine-specific.

The SDK returns raw interleaved PCM rather than a WAV file for both engines. Use `bitsPerSample` when calculating sample counts or constructing a WAV header so consumers remain correct if the addon changes sample width. The complete examples below include a dependency-free WAV writer.

### With ACE-Step

Pass optional lyrics and musical controls to steer the LM:

```ts
import { audioGen } from "@qvac/sdk";

const run = audioGen({
  modelId,
  caption: "Energetic cumbia with brass stabs and live percussion",
  lyrics: `[verse]
The city wakes beneath the moon

[chorus]
Dance until the morning comes`,
  vocalLanguage: "en",
  bpm: 98,
  keyscale: "A minor",
  timesignature: "4/4",
  duration: 30,
  seed: 42,
});

for await (const progress of run.progressStream) {
  console.log(`${progress.stage}: ${progress.step}/${progress.total}`);
}

const [{ pcm, sampleRate, channels, bitsPerSample }, stats] = await Promise.all(
  [run.audio, run.stats],
);
```

Omit `lyrics` or use `"[Instrumental]"` for instrumental output. Omitted musical controls are inferred from the caption. Set `augmentCaptionWithMetadata: true` to also reinforce the BPM, key, and time-signature hints in the engine's internal conditioning caption (the result metadata keeps your original caption); it defaults to `false`.

`duration` is approximate. ACE-Step rounds the requested length to its latent frame grid, so the generated clip can be shorter or longer than requested. Use `stats.audioDurationMs` or calculate the duration from the returned PCM frame count and `sampleRate`; those values describe the actual output.

#### LM-assisted captions, lyrics, and scoring (ACE-Step only)

Four controls put the LM to work around the caption you pass. All four require the default `taskType: "text2music"` and are rejected on any other task:

| Field                 | What it does                                                                                                                                                                                                                                                                                                               |
| --------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `simpleMode`          | Treats `caption` as a short natural-language query and lets the LM compose the full request — a detailed caption, lyrics, and any metadata you left unset. Options you did set are kept. Leave `lyrics` unset for LM-written vocals, or pass `"[Instrumental]"` for an instrumental. Cannot be combined with `audioCodes`. |
| `rewriteQuery`        | Query Rewriting: the LM rewrites `caption` into a detailed musical description while preserving your lyric content. It takes both caption and lyrics as input, so real `lyrics` are required — `"[Instrumental]"` belongs to Simple Mode. Faithful rewriting needs the 1.7B LM.                                            |
| `generateLrc`         | Aligns the lyrics with the generated audio and returns karaoke-style LRC text in `stats.lrc`, with an alignment confidence in `stats.lyricsScore`. Needs lyrics to align: pass `lyrics` or let Simple Mode write them. Instrumental requests are rejected.                                                                 |
| `computeQualityScore` | Teacher-forces the generated codes back through the LM and reports a weighted `[0, 1]` match against the request in `stats.qualityScore` (caption/lyrics PMI plus metadata recall). Costs extra LM forwards after code generation; made for ranking a batch of takes.                                                      |

`simpleMode` and `rewriteQuery` are mutually exclusive — one writes the lyrics, the other rewrites around lyrics you supply — and the SDK rejects a request that sets both.

```ts
// A one-line query, expanded by the LM, scored and timestamped.
const run = audioGen({
  modelId,
  caption: "a hopeful indie track for a road-trip montage",
  simpleMode: true,
  generateLrc: true,
  computeQualityScore: true,
});

const { lrc, lyricsScore, qualityScore } = (await run.stats) ?? {};
```

#### Output and guidance controls (ACE-Step only)

| Field               | Purpose                                                                                                                                                                                                                                   |
| ------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `normalizeLoudness` | Percentile loudness normalization on the generated audio (default `true`): the 99.999th-percentile sample scales to full scale and the tiny tail above it clips. Set `false` for the raw engine output. Audio edits are never normalized. |
| `guidanceScale`     | DiT classifier-free guidance scale. `0` (the default) resolves automatically — `1.0` on turbo variants, which disables CFG, and `7.0` on base/SFT. Values above `1` run CFG via APG and double the DiT cost per step.                     |

#### Sampling and DCW controls (ACE-Step only)

Advanced per-run controls map 1:1 to the ACE-Step engine and are only needed for reproducible quality comparisons; omit them to keep the engine defaults:

| Field           | Purpose                                                                                                                                                                                                         |
| --------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `lmTemperature` | LM sampling temperature (default 0.85)                                                                                                                                                                          |
| `lmTopP`        | LM nucleus-sampling probability (default 0.9)                                                                                                                                                                   |
| `lmTopK`        | LM top-k cutoff; `0` disables top-k filtering                                                                                                                                                                   |
| `lmCfgScale`    | Classifier-free guidance scale used by the LM                                                                                                                                                                   |
| `lmPhase1`      | Let the LM infer missing metadata before semantic-code generation                                                                                                                                               |
| `dcwEnabled`    | Apply the official Haar DCW correction during DiT sampling (default: true)                                                                                                                                      |
| `dcwScaler`     | DCW low-frequency correction strength (default 0.05)                                                                                                                                                            |
| `dcwHighScaler` | DCW high-frequency correction strength (default 0.02)                                                                                                                                                           |
| `audioCodes`    | Frozen semantic codes (an `Int32Array` or `number[]`) to synthesize instead of running the LM — for example the codes of an earlier take. Up to `AUDIOGEN_MAX_AUDIO_CODES` (3000, ten minutes at the LM's 5 Hz) |

### With MiniMax-Music3

MiniMax-Music3 exposes a smaller input surface. It rejects ACE-Step's musical controls (`bpm`, `keyscale`, `timesignature`, `vocalLanguage`, `augmentCaptionWithMetadata`), the ACE-Step-only LM and DCW knobs and `audioCodes` above, the LM-assisted and output controls (`simpleMode`, `rewriteQuery`, `generateLrc`, `computeQualityScore`, `normalizeLoudness`, `guidanceScale`), every reference/cover field, and both `audioEdit()` and `audioUnderstand()`. Pass the caption and, when useful, any of these MiniMax-only controls:

| Field            | When to set it                                                                                                                                                        |
| ---------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `lyrics`         | Vocal lyrics; use `"[Instrumental]"` for instrumental output.                                                                                                         |
| `seed`           | RNG seed for reproducible generation.                                                                                                                                 |
| `duration`       | Target length in seconds. Internally converted to the model's 25 semantic frames per second. Cannot be combined with `maxFrames`.                                     |
| `maxFrames`      | Explicit semantic-frame cap when you want direct control over length instead of a seconds-based target. Cannot be combined with `duration`.                           |
| `inferenceSteps` | Per-run flow sampling steps; leave unset (or `0`) to use the engine default (20). Raise it to trade generation time for higher quality.                               |
| `cfgScale`       | Per-run flow classifier-free guidance scale; leave unset (or `0`) to use the model default. Raise it to follow the caption more literally, lower it for more freedom. |

```ts
import { audioGen } from "@qvac/sdk";

const run = audioGen({
  modelId,
  caption: "Warm cinematic piano with gentle strings",
  lyrics: "[Instrumental]",
  seed: 7,
  cfgScale: 1.7,
  inferenceSteps: 8,
});

const [{ pcm, sampleRate, channels, bitsPerSample }, stats] = await Promise.all(
  [run.audio, run.stats],
);
```

## Reference audio and covers

Reference, cover, and per-layer generation are **ACE-Step only** — MiniMax-Music3 rejects `referenceAudio`, `sourceAudio`, `taskType`, and `track`.

`referenceAudio` conditions the generated timbre without changing the text-to-music task. `taskType: "cover-nofsq"` re-renders `sourceAudio` with a new caption while keeping its structure, optionally combined with a timbre reference:

```ts
import { audioGen } from "@qvac/sdk";

// Timbre reference only.
const styled = audioGen({
  modelId,
  caption: "slow blues with warm electric guitar",
  lyrics: "[Instrumental]",
  referenceAudio: "/path/to/reference.mp3",
});

// Cover: keep the structure of the source song, change the arrangement.
const cover = audioGen({
  modelId,
  caption: "orchestral arrangement with dramatic strings",
  lyrics: "[Instrumental]",
  taskType: "cover-nofsq",
  sourceAudio: "/path/to/source.wav",
  referenceAudio: "/path/to/reference.wav",
  audioCoverStrength: 1,
  coverNoiseStrength: 0.75,
});
```

* `referenceAudio` — Optional timbre reference. Provide it to nudge the generated timbre toward the reference recording; omit it to let the engine pick freely.
* `taskType` — `"text2music"` (default) for caption-driven generation, `"cover-nofsq"` to re-render `sourceAudio` with a new caption while keeping its structure, or `"lego"` to regenerate a single instrument layer of `sourceAudio`. The vocabulary is exported as `AUDIOGEN_TASK_TYPES`.
* `sourceAudio` — The source recording the task works from. Required when `taskType` is `"cover-nofsq"` or `"lego"`.
* `audioCoverStrength` — Fraction of DiT steps (0..1) that keep the source context. `"cover-nofsq"` currently requires `1`; the SDK rejects any other value, so omit it or pass `1`.
* `coverNoiseStrength` — Blends the initial noise toward the source latent (0..1). Closer to `0` gives more variation; closer to `1` stays nearer the source. Default `0`.
* `track` — The instrument layer `"lego"` regenerates. Required when `taskType` is `"lego"` and rejected on any other task. One of `AUDIOGEN_TRACKS`: `"vocals"`, `"backing_vocals"`, `"drums"`, `"bass"`, `"guitar"`, `"keyboard"`, `"percussion"`, `"strings"`, `"synth"`, `"fx"`, `"brass"`, `"woodwinds"`.

<Callout type="warning">
  `"lego"` runs on the **base** DiT only — the engine rejects both the turbo and
  the sft variants. Every `AUDIOGEN_ACESTEP_*` DiT constant on this page is one
  of those, so a lego request built from them fails in the engine. The base DiT
  is not in the registry variant set yet: pass it as an explicit `ditModelSrc`
  path.
</Callout>

```ts
// Lego: keep the song, rebuild one layer of it.
const newDrums = audioGen({
  modelId,
  caption: "the same song with a busier, live-sounding kit",
  taskType: "lego",
  track: "drums",
  sourceAudio: "/path/to/song.wav",
});
```

Both `referenceAudio` and `sourceAudio` accept either a file path (the SDK decodes `.wav`, `.mp3`, `.m4a`, `.ogg`, `.flac`, and `.aac` server-side; any other extension is read as raw interleaved stereo 48 kHz Float32 LE PCM) or a `Buffer` / `Uint8Array` in that same raw PCM layout (see `AUDIOGEN_INPUT_SAMPLE_RATE` and `AUDIOGEN_INPUT_CHANNELS`). Clips are capped at 600 s.

## Edit an existing recording

`audioEdit()` is **ACE-Step only**. It runs an ordered pipeline of edit operations over one source recording and returns the same run shape as `audioGen()`. Operations execute in array order and may repeat or mix:

* `flow-edit` — re-conditions the whole clip from a `from` prompt (what the source is) to a `to` prompt (what it should become), over an optional `nMin`..`nMax` diffusion window (0..1, defaults 0 and 1) with `nAvg` forward-noise samples per step (default 1). Supported on the turbo DiT variants only (`AUDIOGEN_ACESTEP_V15_TURBO_Q4_K_M`, `AUDIOGEN_ACESTEP_V15_TURBO_Q8_0`); the SFT variant rejects it.
* `repaint` — regenerates the `start`..`end` range (seconds) against a new `caption`, keeping the audio outside the range. Omit `end` to repaint through the end of the source; the range must stay inside the source and span at least one latent frame (1/25 s). `mode` is `"conservative"`, `"balanced"` (default), or `"aggressive"`, and `strength` (0..1, default 0.5) tunes balanced-mode preservation.

```ts
import { audioEdit } from "@qvac/sdk";

const run = audioEdit({
  modelId,
  sourceAudio: "/path/to/song.wav",
  operations: [
    {
      type: "flow-edit",
      from: { caption: "original pop song", lyrics: originalLyrics },
      to: { caption: "guitar pop-rock", lyrics: newLyrics },
    },
    {
      type: "repaint",
      caption: "analog synth solo",
      lyrics: "[Instrumental]",
      start: 10,
      end: 20,
    },
    {
      type: "flow-edit",
      from: { caption: "guitar pop-rock" },
      to: { caption: "dark synthwave" },
    },
  ],
  seed: 22883,
});

for await (const progress of run.progressStream) {
  console.log(`${progress.stage}: ${progress.step}/${progress.total}`);
}
const { pcm, sampleRate, channels, bitsPerSample } = await run.audio;
```

`sourceAudio` takes the same forms as the generation inputs above — a file path decoded server-side, or raw interleaved stereo 48 kHz Float32 LE PCM bytes — and every sample must lie in `[-1, 1]`. Prompt `lyrics` default to `"[Instrumental]"`. `seed` seeds the first operation; each following operation uses `seed + index`. The vocabularies are exported as `AUDIOGEN_EDIT_OPERATIONS` and `AUDIOGEN_REPAINT_MODES`.

## Understand a recording

`audioUnderstand()` is **ACE-Step only**. It runs the generation pipeline backwards: the engine encodes `sourceAudio`, recovers its FSQ semantic codes, and the LM reports a caption and the clip's musical metadata. Because the run produces a description rather than PCM, it returns `description` where `audioGen()` returns `audio`; `requestId`, `progressStream`, `stats`, and `diagnostics` are unchanged.

```ts
import { audioUnderstand } from "@qvac/sdk";

const run = audioUnderstand({
  modelId,
  sourceAudio: "/path/to/song.wav",
  vocalLanguage: "en",
  seed: 11,
});

for await (const progress of run.progressStream) {
  console.log(`${progress.stage}: ${progress.step}/${progress.total}`);
}

const {
  caption,
  bpm,
  keyscale,
  timesignature,
  vocalLanguage,
  duration,
  audioCodes,
} = await run.description;
```

`description` resolves to:

| Field                                               | Meaning                                                                |
| --------------------------------------------------- | ---------------------------------------------------------------------- |
| `caption`                                           | The LM's description of the clip.                                      |
| `bpm`, `keyscale`, `timesignature`, `vocalLanguage` | The musical metadata the LM read off the recording.                    |
| `duration`                                          | The LM's estimate in seconds. The recovered codes fix the true length. |
| `audioCodes`                                        | The recovered FSQ semantic codes, as a plain integer array.            |

`sourceAudio` takes the same forms as every other audio input — a file path decoded server-side, or raw interleaved stereo 48 kHz Float32 LE PCM bytes with every sample in `[-1, 1]` — and is capped at 600 s. `seed`, `vocalLanguage`, `lmTemperature`, `lmTopP`, and `lmTopK` are the only other inputs: `vocalLanguage` is forced into the result instead of the LM's guess, and the rest steer the LM decode exactly as they do on a generation.

The recovered codes are a generation input, so a description round-trips straight back into `audioGen()`:

```ts
const { caption, audioCodes } = await audioUnderstand({ modelId, sourceAudio })
  .description;

// Re-synthesize the same piece without re-running the LM.
const remake = audioGen({ modelId, caption, audioCodes });
```

The same description is repeated on the terminal frame as `stats.understand`, so a caller that only wants stats does not need to read the stream.

## Cancellation

The request ID is available before generation starts, so a stop button can cancel exactly that run:

```ts
import { audioGen, cancel, InferenceCancelledError } from "@qvac/sdk";

const run = audioGen({ modelId, caption: "Ambient electronic music" });
stopButton.onclick = () => cancel({ requestId: run.requestId });

try {
  await run.audio;
} catch (error) {
  if (!(error instanceof InferenceCancelledError)) throw error;
}
```

AudioGen supports hard cancellation on both engines: the addon interrupts the active generation. The stream terminates with `stopReason: "cancelled"`, and the `audio` and `stats` promises reject with `InferenceCancelledError`. `audioEdit()` and `audioUnderstand()` runs cancel the same way, rejecting `audio` and `description` respectively. You can also broad-cancel AudioGen work — generations and edits alike — with `cancel({ modelId, kind: "audiogen" })`.

See [Runtime — Cancellation](/sdk/runtime/cancellation) for cancellation behavior shared across the SDK.

## Example

### ACE-Step

This example loads the registry-hosted ACE-Step models, reports progress, generates an instrumental, and writes the PCM output as WAV.

<Callout type="info">
  The generated Python client exposes the contract-level `audio_gen_stream()`
  method. Unlike the TypeScript `audioGen()` wrapper, it yields progress, base64
  PCM chunks, and the terminal stats frame directly; callers decode and
  concatenate the chunks themselves — see the Python tab.
</Callout>

<Tabs>
  <Tab value="ts" label="TypeScript" default>
    <WrapCode>
      ```ts file=<rootDir>/packages/sdk/examples/audiogen/generate-music.ts title="generate-music.ts" lineNumbers
      import { writeFileSync } from 'node:fs'
      import {
        AUDIOGEN_ACESTEP_5HZ_LM_0_6B_Q8_0,
        AUDIOGEN_ACESTEP_V15_TURBO_Q4_K_M,
        AUDIOGEN_QWEN3_EMBEDDING_0_6B_Q8_0,
        AUDIOGEN_VAE_BF16,
        audioGen,
        loadModel,
        unloadModel,
        type ModelProgressUpdate
      } from '@qvac/sdk'

      // Usage:
      //   bun examples/audiogen/generate-music.ts "lo-fi hip hop, mellow piano" output.wav
      const caption =
        process.argv[2] ?? 'Lo-fi hip hop with mellow piano, soft drums, and a warm bass line'
      const outputPath = process.argv[3] ?? 'audiogen-output.wav'

      let modelId: string | undefined
      const lastLoggedPercentageByDownload = new Map<string, number>()
      const completedDownloads = new Set<string>()

      try {
        console.log('▸ Loading ACE-Step AudioGen models...')
        modelId = await loadModel({
          modelType: 'audiogen',
          modelConfig: {
            textEncModelSrc: AUDIOGEN_QWEN3_EMBEDDING_0_6B_Q8_0,
            lmModelSrc: AUDIOGEN_ACESTEP_5HZ_LM_0_6B_Q8_0,
            ditModelSrc: AUDIOGEN_ACESTEP_V15_TURBO_Q4_K_M,
            vaeModelSrc: AUDIOGEN_VAE_BF16,
            useGPU: true,
            inferenceSteps: 8
          },
          onProgress: (progress: ModelProgressUpdate) => {
            const mb = (bytes: number) => (bytes / 1e6).toFixed(1)
            const label = getDownloadLabel(progress.downloadKey)
            const line =
              `▸ Downloading ${label}: ${progress.percentage.toFixed(0)}% ` +
              `(${mb(progress.downloaded)}/${mb(progress.total)} MB)`

            if (process.stderr.isTTY) {
              process.stderr.write(`\r${line}`)
              if (progress.percentage >= 100 && !completedDownloads.has(progress.downloadKey)) {
                completedDownloads.add(progress.downloadKey)
                process.stderr.write('\n')
              }
              return
            }

            const percentageBucket = Math.floor(progress.percentage / 5) * 5
            const lastLogged = lastLoggedPercentageByDownload.get(progress.downloadKey)
            const isNewCompletion = progress.percentage >= 100 && lastLogged !== 100
            if (lastLogged === undefined || percentageBucket > lastLogged || isNewCompletion) {
              lastLoggedPercentageByDownload.set(progress.downloadKey, percentageBucket)
              process.stderr.write(`${line}\n`)
            }
          }
        })

        console.log(`▸ Model loaded: ${modelId}`)
        console.log(`▸ Generating: ${caption}`)

        const run = audioGen({
          modelId,
          caption,
          lyrics: '[Instrumental]',
          seed: 42,
          duration: 10
        })
        console.log(`▸ requestId: ${run.requestId}`)

        for await (const progress of run.progressStream) {
          const value =
            progress.total > 0 ? `${progress.step}/${progress.total}` : `${progress.step} (indeterminate)`
          console.log(`▸ ${progress.stage}: ${value}`)
        }

        const [audio, stats] = await Promise.all([run.audio, run.stats])
        const wav = createWav(audio.pcm, audio.sampleRate, audio.channels, audio.bitsPerSample)
        writeFileSync(outputPath, wav)

        const samplesPerChannel = audio.pcm.byteLength / (audio.bitsPerSample / 8) / audio.channels
        console.log(
          `▸ Generated ${samplesPerChannel} samples per channel at ` +
            `${audio.sampleRate} Hz (${audio.channels} channels)`
        )
        if (stats) console.log(`▸ Stats: ${JSON.stringify(stats)}`)
        console.log(`▸ Saved ${outputPath}`)

        await unloadModel({ modelId })
        modelId = undefined
        console.log('▸ Model unloaded')
        process.exit(0)
      } catch (error) {
        if (modelId !== undefined) {
          try {
            await unloadModel({ modelId })
          } catch {
            // Preserve the generation error as the primary failure.
          }
        }
        console.error('✖', error)
        process.exit(1)
      }

      function getDownloadLabel(downloadKey: string) {
        return downloadKey.split('/').pop() ?? downloadKey
      }

      function createWav(pcm: Uint8Array, sampleRate: number, channels: number, bitsPerSample: number) {
        const header = new ArrayBuffer(44)
        const view = new DataView(header)
        const blockAlign = channels * (bitsPerSample / 8)

        writeAscii(view, 0, 'RIFF')
        view.setUint32(4, 36 + pcm.byteLength, true)
        writeAscii(view, 8, 'WAVE')
        writeAscii(view, 12, 'fmt ')
        view.setUint32(16, 16, true)
        view.setUint16(20, 1, true)
        view.setUint16(22, channels, true)
        view.setUint32(24, sampleRate, true)
        view.setUint32(28, sampleRate * blockAlign, true)
        view.setUint16(32, blockAlign, true)
        view.setUint16(34, bitsPerSample, true)
        writeAscii(view, 36, 'data')
        view.setUint32(40, pcm.byteLength, true)

        const wav = new Uint8Array(44 + pcm.byteLength)
        wav.set(new Uint8Array(header))
        wav.set(pcm, 44)
        return wav
      }

      function writeAscii(view: DataView, offset: number, value: string) {
        for (let index = 0; index < value.length; index++) {
          view.setUint8(offset + index, value.charCodeAt(index))
        }
      }
      ```
    </WrapCode>
  </Tab>

  <Tab value="python" label="Python">
    <WrapCode>
      ```python file=<rootDir>/packages/sdk-python/examples/audiogen.py title="audiogen.py" lineNumbers
      """Python port of packages/sdk/examples/audiogen/generate-music.ts.

      Generate music with ACE-Step AudioGen. Python exposes the contract-level
      `audio_gen_stream` method, so this example decodes and joins its base64 PCM
      chunks before writing a WAV file.

      RUN:
        python examples/audiogen.py "lo-fi hip hop, mellow piano" audiogen-output.wav
      """

      from __future__ import annotations

      import argparse
      import asyncio
      import base64
      import sys
      import wave

      from tetherto.qvac_sdk import (
          Client,
          generate_client_request_id,
          load_model,
          unload_model,
      )
      from tetherto.qvac_sdk.methods import audio_gen_stream
      from tetherto.qvac_sdk.models import (
          AUDIOGEN_ACESTEP_5HZ_LM_0_6B_Q8_0,
          AUDIOGEN_ACESTEP_V15_TURBO_Q4_K_M,
          AUDIOGEN_QWEN3_EMBEDDING_0_6B_Q8_0,
          AUDIOGEN_VAE_BF16,
      )
      from tetherto.qvac_sdk.schemas import AudioGenStreamRequest


      def print_progress(p) -> None:
          """Print model download progress; pass as `on_progress=` to `load_model`."""
          line = (
              f"▸ Downloading {p.percentage:.0f}% "
              f"({p.downloaded / 1e6:.1f}/{p.total / 1e6:.1f} MB)"
          )
          print(line, end="\r" if sys.stderr.isatty() else "\n", file=sys.stderr)
          if p.percentage >= 100:
              print(file=sys.stderr)


      def write_wav(
          pcm: bytes, sample_rate: int, channels: int, bits_per_sample: int, path: str
      ) -> None:
          with wave.open(path, "wb") as output:
              output.setnchannels(channels)
              output.setsampwidth(bits_per_sample // 8)
              output.setframerate(sample_rate)
              output.writeframes(pcm)


      async def main() -> int:
          parser = argparse.ArgumentParser(description="Generate music with AudioGen")
          parser.add_argument(
              "caption",
              nargs="?",
              default="Lo-fi hip hop with mellow piano, soft drums, and warm bass",
          )
          parser.add_argument("output", nargs="?", default="audiogen-output.wav")
          args = parser.parse_args()

          async with Client() as client:
              transport = client.transport
              model_id: str | None = None

              try:
                  print("▸ Loading ACE-Step AudioGen models...")
                  model_id = await load_model(
                      transport,
                      model_type="audiogen-ggml",
                      model_config={
                          "textEncModelSrc": AUDIOGEN_QWEN3_EMBEDDING_0_6B_Q8_0.src,
                          "lmModelSrc": AUDIOGEN_ACESTEP_5HZ_LM_0_6B_Q8_0.src,
                          "ditModelSrc": AUDIOGEN_ACESTEP_V15_TURBO_Q4_K_M.src,
                          "vaeModelSrc": AUDIOGEN_VAE_BF16.src,
                          "useGPU": True,
                          "inferenceSteps": 8,
                      },
                      on_progress=print_progress,
                  )
                  print(f"▸ Model loaded: {model_id}")

                  request_id = generate_client_request_id()
                  request = AudioGenStreamRequest(
                      model_id=model_id,
                      request_id=request_id,
                      caption=args.caption,
                      lyrics="[Instrumental]",
                      seed=42,
                      duration=10,
                  )

                  print(f"▸ requestId: {request_id}")
                  print(f"▸ Generating: {args.caption}")

                  pcm_chunks: list[bytes] = []
                  sample_rate: int | None = None
                  channels: int | None = None
                  bits_per_sample: int | None = None
                  stats = None

                  async for response in audio_gen_stream(transport, request):
                      if response.progress is not None:
                          progress = response.progress
                          value = (
                              f"{progress.step}/{progress.total}"
                              if progress.total > 0
                              else f"{progress.step} (indeterminate)"
                          )
                          print(f"▸ {progress.stage}: {value}")
                      if response.data is not None:
                          pcm_chunks.append(base64.b64decode(response.data))
                          sample_rate = response.sample_rate
                          channels = response.channels
                          bits_per_sample = response.bits_per_sample
                      if response.done:
                          stats = response.stats
                          if (
                              response.stop_reason is not None
                              and response.stop_reason.value == "cancelled"
                          ):
                              print("▸ Generation cancelled", file=sys.stderr)
                              return 1
                          break

                  if (
                      sample_rate is None
                      or channels is None
                      or bits_per_sample is None
                      or bits_per_sample % 8 != 0
                      or not pcm_chunks
                  ):
                      raise RuntimeError("AudioGen stream ended without audio data")

                  pcm = b"".join(pcm_chunks)
                  write_wav(pcm, sample_rate, channels, bits_per_sample, args.output)
                  samples_per_channel = len(pcm) // (bits_per_sample // 8) // channels
                  print(
                      f"▸ Generated {samples_per_channel} samples per channel at "
                      f"{sample_rate} Hz ({channels} channels)"
                  )
                  if stats is not None:
                      print(f"▸ Stats: {stats.model_dump(by_alias=True)}")
                  print(f"▸ Saved {args.output}")
              except Exception as error:
                  print(f"✖ {error}", file=sys.stderr)
                  return 1
              finally:
                  if model_id is not None:
                      try:
                          await unload_model(transport, model_id)
                          print("▸ Model unloaded")
                      except Exception as unload_error:
                          print(f"✖ Failed to unload model: {unload_error}", file=sys.stderr)

          return 0


      if __name__ == "__main__":
          sys.exit(asyncio.run(main()))
      ```
    </WrapCode>
  </Tab>
</Tabs>

The cover example takes an existing recording (any FFmpeg-decodable file), an optional timbre reference, and a new caption, and writes the `cover-nofsq` result as WAV:

<WrapCode>
  ```ts file=<rootDir>/packages/sdk/examples/audiogen/generate-cover.ts title="generate-cover.ts" lineNumbers
  import { writeFileSync } from 'node:fs'
  import {
    AUDIOGEN_ACESTEP_5HZ_LM_0_6B_Q8_0,
    AUDIOGEN_ACESTEP_V15_TURBO_Q4_K_M,
    AUDIOGEN_QWEN3_EMBEDDING_0_6B_Q8_0,
    AUDIOGEN_VAE_BF16,
    audioGen,
    loadModel,
    unloadModel
  } from '@qvac/sdk'

  // Re-render an existing song with a new caption (ACE-Step "cover-nofsq"),
  // optionally conditioning the timbre on a second reference clip.
  //
  // Usage:
  //   bun examples/audiogen/generate-cover.ts <source.wav|mp3|...> "orchestral arrangement" [reference.wav] [output.wav]
  //
  // Both audio inputs are file paths: the SDK decodes them (any FFmpeg-decodable
  // format) to the 48 kHz stereo float PCM the engine expects. Pass raw
  // interleaved stereo 48 kHz Float32 LE PCM as a Buffer instead when the audio
  // is already in memory.
  const sourcePath = process.argv[2]
  const caption = process.argv[3] ?? 'Orchestral arrangement with dramatic strings'
  const referencePath = process.argv[4]
  const outputPath = process.argv[5] ?? 'audiogen-cover.wav'

  if (!sourcePath) {
    console.error(
      'Usage: bun examples/audiogen/generate-cover.ts <source-audio> "<caption>" [reference-audio] [output.wav]'
    )
    process.exit(1)
  }

  let modelId: string | undefined

  try {
    console.log('▸ Loading ACE-Step AudioGen models...')
    modelId = await loadModel({
      modelType: 'audiogen',
      modelConfig: {
        textEncModelSrc: AUDIOGEN_QWEN3_EMBEDDING_0_6B_Q8_0,
        lmModelSrc: AUDIOGEN_ACESTEP_5HZ_LM_0_6B_Q8_0,
        ditModelSrc: AUDIOGEN_ACESTEP_V15_TURBO_Q4_K_M,
        vaeModelSrc: AUDIOGEN_VAE_BF16,
        useGPU: true
      }
    })
    console.log(`▸ Model loaded: ${modelId}`)
    console.log(`▸ Covering ${sourcePath} as: ${caption}`)

    const run = audioGen({
      modelId,
      caption,
      lyrics: '[Instrumental]',
      taskType: 'cover-nofsq',
      sourceAudio: sourcePath,
      ...(referencePath !== undefined && { referenceAudio: referencePath }),
      // cover-nofsq keeps the full source context; coverNoiseStrength blends the
      // initial noise toward the source latent (0 = pure noise, 1 ≈ source).
      audioCoverStrength: 1,
      coverNoiseStrength: 0.75,
      seed: 22886
    })
    console.log(`▸ requestId: ${run.requestId}`)

    for await (const progress of run.progressStream) {
      const value =
        progress.total > 0 ? `${progress.step}/${progress.total}` : `${progress.step} (indeterminate)`
      console.log(`▸ ${progress.stage}: ${value}`)
    }

    const [audio, stats] = await Promise.all([run.audio, run.stats])
    writeFileSync(
      outputPath,
      createWav(audio.pcm, audio.sampleRate, audio.channels, audio.bitsPerSample)
    )
    if (stats) console.log(`▸ Stats: ${JSON.stringify(stats)}`)
    console.log(`▸ Saved ${outputPath}`)

    await unloadModel({ modelId })
    modelId = undefined
    process.exit(0)
  } catch (error) {
    if (modelId !== undefined) {
      try {
        await unloadModel({ modelId })
      } catch {
        // Preserve the generation error as the primary failure.
      }
    }
    console.error('✖', error)
    process.exit(1)
  }

  function createWav(pcm: Uint8Array, sampleRate: number, channels: number, bitsPerSample: number) {
    const header = new ArrayBuffer(44)
    const view = new DataView(header)
    const blockAlign = channels * (bitsPerSample / 8)

    writeAscii(view, 0, 'RIFF')
    view.setUint32(4, 36 + pcm.byteLength, true)
    writeAscii(view, 8, 'WAVE')
    writeAscii(view, 12, 'fmt ')
    view.setUint32(16, 16, true)
    view.setUint16(20, 1, true)
    view.setUint16(22, channels, true)
    view.setUint32(24, sampleRate, true)
    view.setUint32(28, sampleRate * blockAlign, true)
    view.setUint16(32, blockAlign, true)
    view.setUint16(34, bitsPerSample, true)
    writeAscii(view, 36, 'data')
    view.setUint32(40, pcm.byteLength, true)

    const wav = new Uint8Array(44 + pcm.byteLength)
    wav.set(new Uint8Array(header))
    wav.set(pcm, 44)
    return wav
  }

  function writeAscii(view: DataView, offset: number, value: string) {
    for (let index = 0; index < value.length; index++) {
      view.setUint8(offset + index, value.charCodeAt(index))
    }
  }
  ```
</WrapCode>

The edit example takes an existing recording plus a source and a target caption, runs a Flow-Edit followed by a Repaint of seconds 10–20, and writes the result as WAV:

<WrapCode>
  ```ts file=<rootDir>/packages/sdk/examples/audiogen/edit-music.ts title="edit-music.ts" lineNumbers
  import { writeFileSync } from 'node:fs'
  import {
    AUDIOGEN_ACESTEP_5HZ_LM_0_6B_Q8_0,
    AUDIOGEN_ACESTEP_V15_TURBO_Q4_K_M,
    AUDIOGEN_QWEN3_EMBEDDING_0_6B_Q8_0,
    AUDIOGEN_VAE_BF16,
    audioEdit,
    loadModel,
    unloadModel
  } from '@qvac/sdk'

  // Edit an existing recording with ACE-Step: a Flow-Edit that re-conditions the
  // whole clip from a source prompt to a target prompt, then a Repaint that
  // regenerates one time range against a new prompt. Operations run in order.
  //
  // Usage:
  //   bun examples/audiogen/edit-music.ts <source.wav|mp3|...> "original pop song" "guitar pop-rock" [output.wav]
  //
  // The source must run at least REPAINT_END seconds: the Repaint below asks for
  // a fixed window, and a range past the end of the source is rejected with
  // `repaint.start must be within the source duration`.
  //
  // The source is a file path: the SDK decodes it (any FFmpeg-decodable format)
  // to the 48 kHz stereo float PCM the engine expects. Pass raw interleaved
  // stereo 48 kHz Float32 LE PCM in [-1, 1] as a Buffer instead when the audio
  // is already in memory — for example the `pcm` of an earlier `audioGen()` run,
  // converted from Int16 to Float32.
  // Repaint window, in seconds from the start of the source.
  const REPAINT_START = 10
  const REPAINT_END = 20

  const sourcePath = process.argv[2]
  const fromCaption = process.argv[3] ?? 'Original pop song'
  const toCaption = process.argv[4] ?? 'Guitar pop-rock'
  const outputPath = process.argv[5] ?? 'audiogen-edit.wav'

  if (!sourcePath) {
    console.error(
      'Usage: bun examples/audiogen/edit-music.ts <source-audio> "<from caption>" "<to caption>" [output.wav]'
    )
    process.exit(1)
  }

  let modelId: string | undefined

  try {
    console.log('▸ Loading ACE-Step AudioGen models...')
    modelId = await loadModel({
      modelType: 'audiogen',
      modelConfig: {
        textEncModelSrc: AUDIOGEN_QWEN3_EMBEDDING_0_6B_Q8_0,
        lmModelSrc: AUDIOGEN_ACESTEP_5HZ_LM_0_6B_Q8_0,
        // Flow-Edit needs a turbo DiT variant; Repaint works on every variant.
        ditModelSrc: AUDIOGEN_ACESTEP_V15_TURBO_Q4_K_M,
        vaeModelSrc: AUDIOGEN_VAE_BF16,
        useGPU: true
      }
    })
    console.log(`▸ Model loaded: ${modelId}`)
    console.log(`▸ Editing ${sourcePath}: "${fromCaption}" -> "${toCaption}"`)

    const run = audioEdit({
      modelId,
      sourceAudio: sourcePath,
      operations: [
        {
          type: 'flow-edit',
          from: { caption: fromCaption },
          to: { caption: toCaption }
        },
        {
          // Regenerate the repaint window as a synth solo; the rest is kept.
          type: 'repaint',
          caption: 'analog synth solo',
          lyrics: '[Instrumental]',
          start: REPAINT_START,
          end: REPAINT_END,
          mode: 'balanced',
          strength: 0.5
        }
      ],
      // Seeds the first operation; each following operation uses seed + index.
      seed: 22883
    })
    console.log(`▸ requestId: ${run.requestId}`)

    for await (const progress of run.progressStream) {
      const value =
        progress.total > 0 ? `${progress.step}/${progress.total}` : `${progress.step} (indeterminate)`
      console.log(`▸ ${progress.stage}: ${value}`)
    }

    const [audio, stats] = await Promise.all([run.audio, run.stats])
    writeFileSync(
      outputPath,
      createWav(audio.pcm, audio.sampleRate, audio.channels, audio.bitsPerSample)
    )
    if (stats) console.log(`▸ Stats: ${JSON.stringify(stats)}`)
    console.log(`▸ Saved ${outputPath}`)

    await unloadModel({ modelId })
    modelId = undefined
    process.exit(0)
  } catch (error) {
    if (modelId !== undefined) {
      try {
        await unloadModel({ modelId })
      } catch {
        // Preserve the edit error as the primary failure.
      }
    }
    console.error('✖', error)
    process.exit(1)
  }

  function createWav(pcm: Uint8Array, sampleRate: number, channels: number, bitsPerSample: number) {
    const header = new ArrayBuffer(44)
    const view = new DataView(header)
    const blockAlign = channels * (bitsPerSample / 8)

    writeAscii(view, 0, 'RIFF')
    view.setUint32(4, 36 + pcm.byteLength, true)
    writeAscii(view, 8, 'WAVE')
    writeAscii(view, 12, 'fmt ')
    view.setUint32(16, 16, true)
    view.setUint16(20, 1, true)
    view.setUint16(22, channels, true)
    view.setUint32(24, sampleRate, true)
    view.setUint32(28, sampleRate * blockAlign, true)
    view.setUint16(32, blockAlign, true)
    view.setUint16(34, bitsPerSample, true)
    writeAscii(view, 36, 'data')
    view.setUint32(40, pcm.byteLength, true)

    const wav = new Uint8Array(44 + pcm.byteLength)
    wav.set(new Uint8Array(header))
    wav.set(pcm, 44)
    return wav
  }

  function writeAscii(view: DataView, offset: number, value: string) {
    for (let index = 0; index < value.length; index++) {
      view.setUint8(offset + index, value.charCodeAt(index))
    }
  }
  ```
</WrapCode>

The understand example takes an existing recording and prints the caption, musical metadata, and recovered code count the LM reports for it:

<WrapCode>
  ```ts file=<rootDir>/packages/sdk/examples/audiogen/understand-music.ts title="understand-music.ts" lineNumbers
  import {
    AUDIOGEN_ACESTEP_5HZ_LM_0_6B_Q8_0,
    AUDIOGEN_ACESTEP_V15_TURBO_Q4_K_M,
    AUDIOGEN_QWEN3_EMBEDDING_0_6B_Q8_0,
    AUDIOGEN_VAE_BF16,
    audioUnderstand,
    loadModel,
    unloadModel
  } from '@qvac/sdk'

  // Describe an existing recording with ACE-Step's reverse pipeline: the engine
  // encodes the audio, recovers its FSQ semantic codes, and the LM reports a
  // caption and the clip's musical metadata.
  //
  // Usage:
  //   bun examples/audiogen/understand-music.ts <source.wav|mp3|...> [language hint]
  //
  // The source is a file path: the SDK decodes it (any FFmpeg-decodable format)
  // to the 48 kHz stereo float PCM the engine expects. Pass raw interleaved
  // stereo 48 kHz Float32 LE PCM in [-1, 1] as a Buffer instead when the audio
  // is already in memory.
  const sourcePath = process.argv[2]
  const vocalLanguage = process.argv[3]

  if (!sourcePath) {
    console.error('Usage: bun examples/audiogen/understand-music.ts <source-audio> [language]')
    process.exit(1)
  }

  let modelId: string | undefined

  try {
    console.log('▸ Loading ACE-Step AudioGen models...')
    modelId = await loadModel({
      modelType: 'audiogen',
      modelConfig: {
        textEncModelSrc: AUDIOGEN_QWEN3_EMBEDDING_0_6B_Q8_0,
        lmModelSrc: AUDIOGEN_ACESTEP_5HZ_LM_0_6B_Q8_0,
        ditModelSrc: AUDIOGEN_ACESTEP_V15_TURBO_Q4_K_M,
        vaeModelSrc: AUDIOGEN_VAE_BF16,
        useGPU: true
      }
    })
    console.log(`▸ Model loaded: ${modelId}`)
    console.log(`▸ Analysing ${sourcePath}`)

    const run = audioUnderstand({
      modelId,
      sourceAudio: sourcePath,
      // Forced into the result instead of the LM's guess; omit to let it decide.
      ...(vocalLanguage !== undefined && { vocalLanguage }),
      seed: 11
    })
    console.log(`▸ requestId: ${run.requestId}`)

    for await (const progress of run.progressStream) {
      const value =
        progress.total > 0 ? `${progress.step}/${progress.total}` : `${progress.step} (indeterminate)`
      console.log(`▸ ${progress.stage}: ${value}`)
    }

    const [description, stats] = await Promise.all([run.description, run.stats])
    console.log(`▸ Caption: ${description.caption}`)
    console.log(`▸ ${description.bpm} BPM, ${description.keyscale}, ${description.timesignature}`)
    console.log(`▸ Vocal language: ${description.vocalLanguage}`)
    // The LM's duration is an estimate; the recovered codes fix the true length.
    console.log(`▸ Estimated duration: ${description.duration.toFixed(1)} s`)
    console.log(`▸ Recovered ${description.audioCodes.length} semantic codes`)
    if (stats) console.log(`▸ Stats: ${JSON.stringify({ ...stats, understand: undefined })}`)

    // `audioCodes` feeds straight back into audioGen() to re-synthesize the piece
    // without re-running the LM:
    //   audioGen({ modelId, caption: description.caption, audioCodes: description.audioCodes })

    await unloadModel({ modelId })
    modelId = undefined
    process.exit(0)
  } catch (error) {
    if (modelId !== undefined) {
      try {
        await unloadModel({ modelId })
      } catch {
        // Preserve the understand error as the primary failure.
      }
    }
    console.error('✖', error)
    process.exit(1)
  }
  ```
</WrapCode>

### MiniMax-Music3

This example loads a local MiniMax-Music3 pair from the `AUDIOGEN_MINIMAX_LM_MODEL` and `AUDIOGEN_MINIMAX_SYNTH_MODEL` environment variables, generates an instrumental, and writes the PCM output as WAV. Set `AUDIOGEN_MAX_FRAMES`, `AUDIOGEN_STEPS`, and `AUDIOGEN_CFG_SCALE` to override the MiniMax-only per-run controls (`maxFrames`, `inferenceSteps`, `cfgScale`).

<WrapCode>
  ```ts file=<rootDir>/packages/sdk/examples/audiogen/generate-music-minimax.ts title="generate-music-minimax.ts" lineNumbers
  import { writeFileSync } from 'node:fs'
  import { audioGen, loadModel, unloadModel, type ModelProgressUpdate } from '@qvac/sdk'

  // Usage:
  //   AUDIOGEN_MINIMAX_LM_MODEL=/models/mm3-lm-q8.gguf \
  //   AUDIOGEN_MINIMAX_SYNTH_MODEL=/models/mm3-synth-q8.gguf \
  //   bun run examples/audiogen/generate-music-minimax.ts "warm cinematic piano" output.wav
  const caption = process.argv[2] ?? 'Warm cinematic piano with gentle strings'
  const outputPath = process.argv[3] ?? 'minimax-music3-output.wav'

  let modelId: string | undefined

  try {
    const threads = numberFromEnv('AUDIOGEN_THREADS')
    const maxFrames = numberFromEnv('AUDIOGEN_MAX_FRAMES')
    const inferenceSteps = numberFromEnv('AUDIOGEN_STEPS')

    console.log('▸ Loading MiniMax-Music3 models...')
    modelId = await loadModel({
      modelType: 'audiogen',
      modelConfig: {
        engine: 'minimax',
        lmModelSrc: requiredModelPath('AUDIOGEN_MINIMAX_LM_MODEL'),
        synthModelSrc: requiredModelPath('AUDIOGEN_MINIMAX_SYNTH_MODEL'),
        useGPU: process.env['AUDIOGEN_USE_GPU'] === '1',
        ...(threads !== undefined && { threads })
      },
      onProgress: function (progress: ModelProgressUpdate) {
        logDownloadProgress(progress)
      }
    })

    console.log(`▸ Model loaded: ${modelId}`)
    console.log(`▸ Generating: ${caption}`)
    const run = audioGen({
      modelId,
      caption,
      lyrics: process.env['AUDIOGEN_LYRICS'] ?? '[Instrumental]',
      seed: numberFromEnv('AUDIOGEN_SEED') ?? 7,
      cfgScale: numberFromEnv('AUDIOGEN_CFG_SCALE') ?? 1.7,
      ...(maxFrames !== undefined && { maxFrames }),
      ...(inferenceSteps !== undefined && { inferenceSteps })
    })

    console.log(`▸ requestId: ${run.requestId}`)
    for await (const progress of run.progressStream) {
      const value =
        progress.total > 0 ? `${progress.step}/${progress.total}` : `${progress.step} (indeterminate)`
      console.log(`▸ ${progress.stage}: ${value}`)
    }

    const [audio, stats] = await Promise.all([run.audio, run.stats])
    writeFileSync(
      outputPath,
      createWav(audio.pcm, audio.sampleRate, audio.channels, audio.bitsPerSample)
    )
    if (stats) console.log(`▸ Stats: ${JSON.stringify(stats)}`)
    console.log(`▸ Saved ${outputPath}`)

    await unloadModel({ modelId })
    modelId = undefined
    console.log('▸ Model unloaded')
    process.exit(0)
  } catch (error) {
    if (modelId !== undefined) {
      try {
        await unloadModel({ modelId })
      } catch {
        // Preserve the generation error as the primary failure.
      }
    }
    console.error('✖', error)
    process.exit(1)
  }

  function requiredModelPath(name: string) {
    const value = process.env[name]
    if (!value) throw new Error(`${name} is required`)
    return value
  }

  function numberFromEnv(name: string) {
    const value = process.env[name]
    if (value === undefined) return undefined
    if (value.trim() === '') throw new Error(`${name} must not be empty`)

    const number = Number(value)
    if (!Number.isFinite(number)) throw new Error(`${name} must be a finite number`)
    return number
  }

  function logDownloadProgress(progress: ModelProgressUpdate) {
    const downloadedMb = (progress.downloaded / 1e6).toFixed(1)
    const totalMb = (progress.total / 1e6).toFixed(1)
    const line = `▸ Downloading ${progress.percentage.toFixed(0)}% (${downloadedMb}/${totalMb} MB)`
    process.stderr.write(process.stderr.isTTY ? `\r${line}` : `${line}\n`)
    if (process.stderr.isTTY && progress.percentage >= 100) process.stderr.write('\n')
  }

  function createWav(pcm: Uint8Array, sampleRate: number, channels: number, bitsPerSample: number) {
    const header = new ArrayBuffer(44)
    const view = new DataView(header)
    const blockAlign = channels * (bitsPerSample / 8)

    writeAscii(view, 0, 'RIFF')
    view.setUint32(4, 36 + pcm.byteLength, true)
    writeAscii(view, 8, 'WAVE')
    writeAscii(view, 12, 'fmt ')
    view.setUint32(16, 16, true)
    view.setUint16(20, 1, true)
    view.setUint16(22, channels, true)
    view.setUint32(24, sampleRate, true)
    view.setUint32(28, sampleRate * blockAlign, true)
    view.setUint16(32, blockAlign, true)
    view.setUint16(34, bitsPerSample, true)
    writeAscii(view, 36, 'data')
    view.setUint32(40, pcm.byteLength, true)

    const wav = new Uint8Array(44 + pcm.byteLength)
    wav.set(new Uint8Array(header))
    wav.set(pcm, 44)
    return wav
  }

  function writeAscii(view: DataView, offset: number, value: string) {
    for (let index = 0; index < value.length; index++) {
      view.setUint8(offset + index, value.charCodeAt(index))
    }
  }
  ```
</WrapCode>

<Callout type="success">
  **Tip:** see the [JS/TS quickstart](/sdk/js-ts-sdk#quickstart) or the [Python
  quickstart](/sdk/python-sdk#quickstart) for setup instructions, and apply the
  first-run download configuration above on constrained connections.
</Callout>
