New: TranslatePsy-AfriSLM translates directly between 19 African languages, offline.
QVAC Logo

@qvac/audiogen-ggml

Native music generation with ACE-Step 1.5 and GGML.

Overview

@qvac/audiogen-ggml is a Bare native addon for generating music from captions, lyrics, and musical controls, optionally conditioned on a reference recording or re-rendering an existing song as a cover. It runs ACE-Step 1.5 locally with CPU inference or optional Metal/Vulkan acceleration (desktop, Android, and iOS prebuilds) and returns interleaved stereo Int16 PCM.

SDK consumers should use audioGen() through @qvac/sdk. Use this package directly only when building against the addon-level Bare API.

Pipeline

Generation runs through four persistent model stages:

  1. A Qwen3 text encoder for captions and lyrics.
  2. An ACE-Step language model for song structure and musical controls.
  3. A DiT model for latent audio generation.
  4. A VAE decoder for the final waveform.

The models are loaded once and reused between runs.

Models

StageFile
Text encoderQwen3-Embedding-0.6B-Q8_0.gguf
Language modelacestep-5Hz-lm-0.6B-Q8_0.gguf
Turbo DiT, Q4acestep-v15-turbo-Q4_K_M.gguf
Turbo DiT, Q8acestep-v15-turbo-Q8_0.gguf
SFT DiT, Q8acestep-v15-sft-Q8_0.gguf
VAEvae-BF16.gguf

Choose one DiT variant and combine it with the fixed text encoder, language model, and VAE.

Installation

npm install @qvac/audiogen-ggml

The package requires a supported native prebuild and the four GGUF files. Models are not bundled with the addon.

Minimal Bare usage

const { AudioGen } = require("@qvac/audiogen-ggml");

const model = new AudioGen({
  files: {
    modelDir: "/path/to/acestep/models",
    ditVariant: "turbo-q4",
  },
  config: { useGPU: true },
});

await model.load();

const response = await model.run("lo-fi hip hop with mellow piano", {
  lyrics: "[Instrumental]",
  seed: 42,
});

for await (const update of response.iterate()) {
  if (update.outputArray) {
    // Interleaved Int16 PCM at update.sampleRate.
  }
}

const stats = await response.await();
await model.destroy();

destroy() is terminal: create a new AudioGen instance instead of calling load() again on a destroyed one.

Reference and cover audio

run() also accepts referenceAudio (timbre conditioning) and, with taskType: "cover-nofsq", sourceAudio plus audioCoverStrength / coverNoiseStrength. Both PCM inputs are Float32Array values holding finite, normalized samples in interleaved stereo order at 48 kHz; the addon does not resample or convert channels. Through the SDK, audioGen() accepts file paths (decoded for you) or raw PCM buffers — see Music generation. The addon's ordered edit() pipeline (Flow-Edit and Repaint) is exposed by the SDK as audioEdit(), and its reverse understand() pipeline as audioUnderstand().

For runtime controls, LM sampling and DCW parameters, supported encoders, cancellation, and complete examples, see the package README.

On this page

Ask anything about QVAC.