---
title: "@qvac/audiogen-ggml"
canonical: https://docs.qvac.tether.io/ecosystem/addons/audiogen-ggml/
collection: "Ecosystem"
---

# @qvac/audiogen-ggml (/ecosystem/addons/audiogen-ggml)



## Overview

[`@qvac/audiogen-ggml`](https://github.com/tetherto/qvac/tree/main/packages/audiogen-ggml) is a Bare native addon for generating music from captions, lyrics, and musical controls, optionally conditioned on a reference recording or re-rendering an existing song as a cover. It runs ACE-Step 1.5 locally with CPU inference or optional Metal/Vulkan acceleration (desktop, Android, and iOS prebuilds) and returns interleaved stereo Int16 PCM.

<Callout type="info">
  SDK consumers should use [`audioGen()` through
  `@qvac/sdk`](/sdk/ai-capabilities/music-generation). Use this package directly
  only when building against the addon-level Bare API.
</Callout>

## Pipeline

Generation runs through four persistent model stages:

1. A Qwen3 text encoder for captions and lyrics.
2. An ACE-Step language model for song structure and musical controls.
3. A DiT model for latent audio generation.
4. A VAE decoder for the final waveform.

The models are loaded once and reused between runs.

## Models

| Stage          | File                             |
| -------------- | -------------------------------- |
| Text encoder   | `Qwen3-Embedding-0.6B-Q8_0.gguf` |
| Language model | `acestep-5Hz-lm-0.6B-Q8_0.gguf`  |
| Turbo DiT, Q4  | `acestep-v15-turbo-Q4_K_M.gguf`  |
| Turbo DiT, Q8  | `acestep-v15-turbo-Q8_0.gguf`    |
| SFT DiT, Q8    | `acestep-v15-sft-Q8_0.gguf`      |
| VAE            | `vae-BF16.gguf`                  |

Choose one DiT variant and combine it with the fixed text encoder, language model, and VAE.

## Installation

```bash
npm install @qvac/audiogen-ggml
```

The package requires a supported native prebuild and the four GGUF files. Models are not bundled with the addon.

## Minimal Bare usage

```js
const { AudioGen } = require("@qvac/audiogen-ggml");

const model = new AudioGen({
  files: {
    modelDir: "/path/to/acestep/models",
    ditVariant: "turbo-q4",
  },
  config: { useGPU: true },
});

await model.load();

const response = await model.run("lo-fi hip hop with mellow piano", {
  lyrics: "[Instrumental]",
  seed: 42,
});

for await (const update of response.iterate()) {
  if (update.outputArray) {
    // Interleaved Int16 PCM at update.sampleRate.
  }
}

const stats = await response.await();
await model.destroy();
```

`destroy()` is terminal: create a new `AudioGen` instance instead of calling `load()` again on a destroyed one.

## Reference and cover audio

`run()` also accepts `referenceAudio` (timbre conditioning) and, with `taskType: "cover-nofsq"`, `sourceAudio` plus `audioCoverStrength` / `coverNoiseStrength`. Both PCM inputs are `Float32Array` values holding finite, normalized samples in interleaved stereo order at 48 kHz; the addon does not resample or convert channels. Through the SDK, `audioGen()` accepts file paths (decoded for you) or raw PCM buffers — see [Music generation](/sdk/ai-capabilities/music-generation#reference-audio-and-covers). The addon's ordered `edit()` pipeline (Flow-Edit and Repaint) is exposed by the SDK as [`audioEdit()`](/sdk/ai-capabilities/music-generation#edit-an-existing-recording), and its reverse `understand()` pipeline as [`audioUnderstand()`](/sdk/ai-capabilities/music-generation#understand-a-recording).

For runtime controls, LM sampling and DCW parameters, supported encoders, cancellation, and complete examples, see the [package README](https://github.com/tetherto/qvac/tree/main/packages/audiogen-ggml#readme).
