---
title: "@qvac/tts-ggml"
canonical: https://docs.qvac.tether.io/ecosystem/addons/tts-ggml/
collection: "Ecosystem"
---

# @qvac/tts-ggml (/ecosystem/addons/tts-ggml)



## Overview

[Bare module](https://bare.pears.com) that adds support for text-to-speech in QVAC, backed by the [`qvac-tts.cpp`](https://github.com/tetherto/qvac-fabric-speech.cpp/tree/master/engines/tts) GGML library.

It runs in-process with a persistent native engine — the GGUFs, the S3Gen preload, the ggml backend, and any voice-conditioning tensors are loaded once and reused across every synthesis call. GPU acceleration (Metal on macOS/iOS, CUDA on NVIDIA Linux x64, Vulkan / OpenCL on Linux/Windows/Android) is **opt-in** via `config: { useGPU: true }`; the default is CPU.

## Models

Six engine families are wrapped — **Chatterbox**, **Supertonic**,
**Parler-TTS**, **CosyVoice3**, **Audio8**, and **MOSS** — each with its own
model layout under `models/`:

| Model                   | GGUF files                                                                                             | Languages / notes                                                                                                                                     |
| ----------------------- | ------------------------------------------------------------------------------------------------------ | ----------------------------------------------------------------------------------------------------------------------------------------------------- |
| Chatterbox Turbo        | `chatterbox-t3-turbo.gguf`, `chatterbox-s3gen.gguf`                                                    | English; voice cloning                                                                                                                                |
| Chatterbox multilingual | `chatterbox-t3-mtl.gguf`, `chatterbox-s3gen-mtl.gguf`                                                  | 23 languages (see below); voice cloning                                                                                                               |
| Supertonic              | `supertonic.gguf`                                                                                      | English; voice baked in                                                                                                                               |
| Supertonic 2            | `supertonic2.gguf`                                                                                     | Multilingual: `en`, `ko`, `es`, `pt`, `fr`                                                                                                            |
| Supertonic 3            | `supertonic3.gguf`                                                                                     | Multilingual: 31 languages plus the language-agnostic `na` code (see below)                                                                           |
| Parler-TTS Mini v1      | `parler-tts-mini-v1.gguf`                                                                              | English; description-conditioned voices                                                                                                               |
| CosyVoice3              | `cosyvoice3/` dir (`cosyvoice3-{llm,flow,hift}-*.gguf` + `voice.gguf` + tokenizer files)               | Instruct-conditioned; strong on Chinese + 17 dialects; native 24 kHz                                                                                  |
| Audio8                  | `audio8-lm-q8_0.gguf` + `audio8-codec-decoder-q8_0.gguf` (+ `audio8-codec-encoder-q8_0.gguf` to clone) | Multilingual (language inferred from the text); voice cloning; native 44.1 kHz                                                                        |
| MOSS-TTS v1.5           | `moss-tts-delay-f16.gguf` + `moss-codec-decoder-f16.gguf` (+ `moss-codec-encoder-f16.gguf` to clone)   | Multilingual (`language` hint); directable pauses, duration and pronunciation; voice cloning; native chunk streaming; native 24 kHz; 8B, desktop only |
| MOSS-TTSD               | `moss-ttsd-f16.gguf` + `moss-codec-decoder-f16.gguf` + `moss-codec-encoder-f16.gguf`                   | Multi-speaker dialogue, one reference recording per speaker; native 24 kHz; 8B, desktop only                                                          |

**Chatterbox multilingual (23 languages):** Arabic (`ar`), Danish (`da`), German (`de`), Greek (`el`), English (`en`), Spanish (`es`), Finnish (`fi`), French (`fr`), Hebrew (`he`), Hindi (`hi`), Italian (`it`), Japanese (`ja`), Korean (`ko`), Malay (`ms`), Dutch (`nl`), Norwegian (`no`), Polish (`pl`), Portuguese (`pt`), Russian (`ru`), Swedish (`sv`), Swahili (`sw`), Turkish (`tr`), Chinese (`zh`).

**Supertonic 3 (31 languages + `na`):** `ar`, `bg`, `hr`, `cs`, `da`, `nl`, `en`, `et`, `fi`, `fr`, `de`, `el`, `hi`, `hu`, `id`, `it`, `ja`, `ko`, `lv`, `lt`, `pl`, `pt`, `ro`, `ru`, `sk`, `sl`, `es`, `sv`, `tr`, `uk`, `vi`. Pass `language: 'na'` when the input language is unknown.

Point the addon at a custom location via `files.modelDir` (engine auto-detected from the GGUF filenames present), or pass explicit `files.t3Model` + `files.s3genModel` (Chatterbox), `files.supertonicModel` (Supertonic), `files.parlerModel` (Parler-TTS), `files.cosyvoiceModelDir` (CosyVoice3), `files.audio8Lm` + `files.audio8CodecDecoder` (+ `files.audio8CodecEncoder` for voice cloning) (Audio8), or `files.mossBackbone` + `files.mossCodecDecoder` (+ `files.mossCodecEncoder` for voice cloning and dialogue) (MOSS). Use `npm run download-models:registry` to acquire registry-published Chatterbox, Supertonic, Parler, and Audio8 models. CosyVoice3 and MOSS must currently be staged from local converted artifacts.

## Requirement

Bare $\geq$ v1.19

## Installation

```bash
npm i @qvac/tts-ggml
```

`@qvac/tts-ggml` is a meta package that ships the JavaScript wrapper only.
The native prebuild for each desktop host installs through a version-locked,
`os`/`cpu` filtered optional dependency:

| Host                | Package                       |
| ------------------- | ----------------------------- |
| linux-x64 (glibc)   | `@qvac/tts-ggml-linux-x64`    |
| linux-arm64 (glibc) | `@qvac/tts-ggml-linux-arm64`  |
| darwin-arm64        | `@qvac/tts-ggml-darwin-arm64` |
| darwin-x64          | `@qvac/tts-ggml-darwin-x64`   |
| win32-x64           | `@qvac/tts-ggml-win32-x64`    |

Do not depend on the desktop platform packages directly. Supported installers
are npm 7+, pnpm, bun, and Yarn Berry; Yarn v1 and `--omit=optional` installs
skip the platform package and fail at require time with an error naming it. A
locally built `prebuilds/` directory in the package root always takes
precedence, and `require('@qvac/tts-ggml').resolveBackendsDir()` locates the
directory holding the host's binaries and dynamically loaded ggml backends.
Unsupported targets require an explicit source build; installation does not
automatically compile a local addon.

Mobile targets are cross-built, so no install host ever matches their `os`,
and optional-dependency filtering can never select them. Mobile applications
must declare the target's platform package as a direct dependency, pinned to
the exact `@qvac/tts-ggml` version:

| Target                    | Package                        |
| ------------------------- | ------------------------------ |
| android-arm64             | `@qvac/tts-ggml-android-arm64` |
| ios (device + simulators) | `@qvac/tts-ggml-ios`           |

```json
{
  "dependencies": {
    "@qvac/tts-ggml": "x.y.z",
    "@qvac/tts-ggml-android-arm64": "x.y.z"
  }
}
```

The Linux x64 prebuild bundles the CUDA and Vulkan backends as runtime-loaded
modules (CUDA is also available as an opt-in `ENABLE_CUDA=ON` source build on
Linux arm64 and Windows x64). With `useGPU: true` the engine prefers CUDA when
the NVIDIA driver and the CUDA 13 runtime libraries (`cudart`, `cuBLAS`)
resolve at load time; otherwise the CUDA module is skipped and it falls back to
Vulkan or CPU. Where more than one backend is usable, `TTS_CPP_GPU_BACKEND`
(`cuda` | `vulkan` | `metal` | `opencl`) pins the choice — the SDK worker
inherits the host process environment, so exporting it before starting the SDK
is enough.

## Quickstart

<Steps>
  <Step>
    If you don't have Bare runtime, install it:

    ```bash
    npm i -g bare
    ```
  </Step>

  <Step>
    Create a new project:

    ```bash
    mkdir qvac-tts-quickstart
    cd qvac-tts-quickstart
    npm init -y
    ```
  </Step>

  <Step>
    Install dependencies:

    ```bash
    npm i @qvac/tts-ggml bare-fs bare-path
    ```
  </Step>

  <Step>
    Place the Chatterbox GGUF files into `models/`: `chatterbox-t3-turbo.gguf` and
    `chatterbox-s3gen.gguf`. Optionally place a mono reference WAV (≥ 5 s of clean
    speech) at `./reference.wav` for voice cloning.
  </Step>

  <Step>
    Create 

    `index.js`

    :
  </Step>

  <WrapCode>
    ```js title="index.js" lineNumbers
    "use strict";

    const fs = require("bare-fs");
    const TTSGgml = require("@qvac/tts-ggml");

    const SAMPLE_RATE = 24000;

    async function main() {
      const model = new TTSGgml({
        files: { modelDir: "./models" }, // contains chatterbox-{t3-turbo,s3gen}.gguf
        referenceAudio: "./reference.wav", // optional voice cloning
        config: { language: "en" },
        opts: { stats: true },
      });

      try {
        console.log("Loading Chatterbox TTS model...");
        await model.load();
        console.log("Model loaded.");

        const textToSynthesize =
          "Hello world! This is a test of the Chatterbox TTS system.";
        console.log(`Running TTS on: "${textToSynthesize}"`);

        const response = await model.run({
          input: textToSynthesize,
          type: "text",
        });

        let pcm = [];
        await response
          .onUpdate((data) => {
            if (data && data.outputArray)
              pcm = pcm.concat(Array.from(data.outputArray));
          })
          .await();

        console.log("TTS finished!");
        if (response.stats) {
          console.log(`Inference stats: ${JSON.stringify(response.stats)}`);
        }

        console.log(`Generated ${pcm.length} audio samples at ${SAMPLE_RATE}Hz`);
      } catch (err) {
        console.error("Error during TTS processing:", err);
      } finally {
        console.log("Unloading model...");
        await model.unload();
        console.log("Model unloaded.");
      }
    }

    main().catch(console.error);
    ```
  </WrapCode>

  <Step>
    Run `index.js`:

    ```bash
    bare index.js
    ```
  </Step>
</Steps>

## Usage

### 1. Import the Model Class

```js
const TTSGgml = require("@qvac/tts-ggml");
```

### 2. Create the Model Instance

```js
const model = new TTSGgml({
  files: { modelDir: "./models" },
  referenceAudio: "./voices/me.wav",
  config: { language: "en", useGPU: false },
  opts: { stats: true },
});
```

The most common constructor options:

| Option                                          | Type             | Default         | Description                                                                                                   |
| ----------------------------------------------- | ---------------- | --------------- | ------------------------------------------------------------------------------------------------------------- |
| `files.modelDir`                                | string           | —               | Directory containing the two GGUFs (engine auto-detected)                                                     |
| `files.t3Model` / `files.s3genModel`            | string           | —               | Override `modelDir` for the Chatterbox T3 / S3Gen GGUF                                                        |
| `files.supertonicModel`                         | string           | —               | Supertonic GGUF path                                                                                          |
| `files.parlerModel`                             | string           | —               | Parler GGUF path (mini / large / indic variant)                                                               |
| `files.cosyvoiceModelDir`                       | string           | —               | CosyVoice3 model directory (`cosyvoice3-{llm,flow,hift}-*.gguf` + `voice.gguf` + tokenizer files)             |
| `files.audio8Lm` / `files.audio8CodecDecoder`   | string           | —               | Audio8 LM / codec decoder GGUFs (override `modelDir`)                                                         |
| `files.audio8CodecEncoder`                      | string           | —               | Audio8 codec encoder — only needed to clone a voice                                                           |
| `files.mossBackbone` / `files.mossCodecDecoder` | string           | —               | MOSS backbone (`moss-tts-delay-*` or `moss-ttsd-*`) / codec decoder GGUFs (override `modelDir`)               |
| `files.mossCodecEncoder`                        | string           | —               | MOSS codec encoder — only needed to clone a voice or for `dialogueReferences`                                 |
| `referenceAudio`                                | string           | —               | Mono WAV for voice cloning (Chatterbox: ≥ 5 s; Audio8: also needs `referenceText`; MOSS: 24 kHz)              |
| `dialogueReferences`                            | string\[]        | —               | MOSS-TTSD: one 24 kHz WAV per speaker (`[S1]`, `[S2]`, …); the text opens with their transcripts              |
| `durationTokens`                                | number           | 0               | MOSS: target length in codec frames (12.5 per second); `0` keeps it free                                      |
| `referenceText`                                 | string           | —               | Audio8-only: what `referenceAudio` says, verbatim; required with a reference                                  |
| `voiceDir`                                      | string           | —               | Pre-baked voice profile directory                                                                             |
| `emotion` / `pace`                              | string           | —               | Cross-engine conditioning — Parler + CosyVoice3 emotions; Parler / CosyVoice3 / Supertonic pace               |
| `instruct`                                      | object \| string | —               | CosyVoice3-only: `dialect` / `volume` / `style` object (precedence in that order) or a raw instruction string |
| `temperature` / `topK` / `topP`                 | number           | engine defaults | Parler + Audio8 sampling knobs (Audio8: temp 0.7 / top-k 50 / top-p 0.9)                                      |
| `greedy`                                        | boolean          | `false`         | Audio8-only: take the argmax; ignores `temperature` / `topK` / `topP`                                         |
| `streamChunkTokens`                             | number           | 0               | `> 0` enables native chunk streaming (25 tokens ≈ 1 s of audio)                                               |
| `cfmSteps`                                      | number           | 2               | `1` halves CFM cost for faster synthesis                                                                      |
| `config.language`                               | string           | `'en'`          | Language code; multilingual models accept `es/fr/de/pt/it/zh/ja/ko/...`                                       |
| `config.useGPU`                                 | boolean          | `false`         | Route through Metal / CUDA / Vulkan / OpenCL if available                                                     |
| `config.outputSampleRate`                       | number           | 24000           | Resample the native 24 kHz output                                                                             |
| `opts.stats`                                    | boolean          | `false`         | Populate `response.stats` with RTF, backend info, etc.                                                        |

See the [package README](https://github.com/tetherto/qvac/tree/main/packages/tts-ggml) for the full option set (GPU/backend, KV-cache, and Android-specific options).

### 3. Load the Model

```js
await model.load();
```

`load()` constructs the native engine — it loads T3, preloads S3Gen, and bakes voice conditioning. Subsequent `run()` calls reuse all of it.

### 4. Run TTS Synthesis

Pass the text to synthesize to the `run` method and process the generated audio output asynchronously:

```javascript
try {
  const textToSynthesize = "Hello world! This is a test of the TTS system.";
  let audioSamples = [];

  const response = await model.run({
    input: textToSynthesize,
    type: "text",
  });

  await response
    .onUpdate((data) => {
      if (data && data.outputArray) {
        audioSamples = audioSamples.concat(Array.from(data.outputArray));
      }
    })
    .await();

  console.log(`Total audio samples generated: ${audioSamples.length}`);

  // audioSamples now contains the complete audio as PCM data (16-bit, 24 kHz, mono)
  if (response.stats) {
    console.log(`Inference stats: ${JSON.stringify(response.stats)}`);
  }
} catch (error) {
  console.error("TTS synthesis failed:", error);
}
```

### 5. Release Resources

Unload the model when finished:

```javascript
try {
  await model.unload();
} catch (error) {
  console.error("Failed to unload model:", error);
}
```

## Streaming

### Sentence streaming — `runStreaming(asyncIterable)`

Use when your text arrives as discrete sentences (e.g. buffered LLM output) and you want the audio to flow sentence-by-sentence. One `onUpdate` event per input yield:

```js
async function* sentencesOverTime() {
  yield "First sentence.";
  await new Promise((r) => setTimeout(r, 200));
  yield "The second arrives shortly after.";
}

const response = await model.runStreaming(sentencesOverTime());
await response
  .onUpdate((data) => {
    // data.outputArray   — Int16 PCM for this sentence's audio
    // data.chunkIndex    — 0-based index of the yielded sentence
    // data.sentenceChunk — the sentence text that produced this audio
  })
  .await();
```

`runStreaming(textStream, options)` accepts a string, string array, iterable,
or async iterable. Async iterables default to `accumulateSentences: true`;
strings, arrays, and synchronous iterables default to one synthesis job per
item. Set `accumulateSentences` explicitly to override that behavior. Select
`sentenceDelimiterPreset: "latin"`, `"multilingual"`, or `"cjk"`, or provide a
`sentenceDelimiter` regular expression. `maxBufferScalars` forces a flush at
the buffer limit and `flushAfterMs` flushes incomplete text after the configured
delay. Parler description fields and Audio8 `referenceAudio` / `referenceText`
can also be fixed for the full streaming response.

### Chunk streaming — `streamChunkTokens`

Use when you want the fastest possible first-audio-out **within a single utterance**. The C++ engine splits each synthesis into chunks of `streamChunkTokens` speech tokens and emits audio per chunk:

```js
const model = new TTSGgml({
  files: { modelDir: "./models" },
  streamChunkTokens: 25, // ~1 s of audio per chunk
  streamFirstChunkTokens: 10, // smaller first chunk = faster first-audio-out
  cfmSteps: 1,
  config: { language: "en" },
});

await model.load();

const response = await model.run({
  input: "A long sentence produces many chunks...",
});
await response
  .onUpdate((data) => {
    if (data && data.outputArray) playPcmChunk(data.outputArray);
  })
  .await();
```

## Voice cloning

Pass a mono WAV with ≥ 5 s of clean speech. The engine does the loudness normalisation, resampling, and all conditioning natively at `load()` time:

```js
const model = new TTSGgml({
  files: { modelDir: "./models" },
  referenceAudio: "./voices/me.wav",
  config: { language: "en" },
});
```

Alternatively point at a pre-baked profile directory via `voiceDir`. When both are supplied, missing tensors in `voiceDir` are backfilled from `referenceAudio`.

**Audio8** clones from the recording plus **what is said in it**: supply
`referenceAudio` and `referenceText` (the verbatim transcript) together with
`files.audio8CodecEncoder`, which encodes the reference to codes in-process —
no enrolment step, no voice profile. The reference can also be switched per
call. **CosyVoice3** does not take a reference — it always speaks with its
baked default voice.

## Speech enhancement (LavaSR)

Opt-in neural post-processing that bandwidth-extends the synthesized audio to **48 kHz** with a synthesised high band, using the LavaSR Vocos enhancer run on the CPU/GGML path. It is fully backward compatible — provide no enhancer GGUF and nothing changes. Enhancement is enabled simply by supplying the enhancer GGUF; there is no separate on/off flag.

```js
const model = new TTSGgml({
  engine: TTSGgml.ENGINE_SUPERTONIC,
  // Providing the enhancer GGUF is what turns enhancement on:
  files: {
    supertonicModel,
    lavasrEnhancer: "models/lavasr/lavasr-enhancer.gguf",
  },
  config: { language: "en" },
});
// The output callback now reports 48000:
//   response.onUpdate(d => )
```

The GGUF path may instead be given as an `enhancer: { type: 'lavasr', enhancerPath }` block.

* Works for Chatterbox, Supertonic, Parler, and CosyVoice3 (Audio8 is not
  supported) — on the batch path, sentence-level streaming, and native chunk
  streaming where the selected engine supports it.
* For native chunk streaming the enhancer runs over a sliding window with look-ahead + crossfade, so each emitted chunk is bandwidth-extended seam-free; this adds \~0.34 s of look-ahead latency.
* The enhancer always runs at 48 kHz internally. By default the emitted audio is 48 kHz; set `config.outputSampleRate` to resample the enhanced output (`sampleRate` reports the actual rate).

### Denoiser

LavaSR's first stage — the UL-UNAS **denoiser** that cleans the signal before the enhancer bandwidth-extends it — is enabled the same way, via `files.lavasrDenoiser` (or a `denoiser: { type: 'lavasr', denoiserPath }` block), and runs before the enhancer on the batch path:

```js
const model = new TTSGgml({
  engine: TTSGgml.ENGINE_SUPERTONIC,
  files: {
    supertonicModel,
    lavasrDenoiser: "models/lavasr/lavasr-denoiser.gguf", // cleaned first…
    lavasrEnhancer: "models/lavasr/lavasr-enhancer.gguf", // …then upsampled
  },
  config: { language: "en" },
});
```

* The denoiser forward runs at 16 kHz internally (resampled in/out), so it is **rate-preserving** — the emitted audio keeps the engine's sample rate. With no denoiser path the output is unchanged.
* Denoiser + Chatterbox native chunk streaming (`streamChunkTokens > 0`) is rejected up front; use batch synthesis, or drop the denoiser for streaming.

## Output Format

Audio is received via the `onUpdate` callback of the response object as raw PCM samples.

```javascript
response.onUpdate((data) => {
  data.outputArray; // Int16Array — mono PCM
  data.sampleRate; // actual output sample rate
  data.chunkIndex; // present on sentence-streaming events only
  data.sentenceChunk; // present on sentence-streaming events only
});
```

When synthesis completes and `opts: { stats: true }` was set, `response.stats` reports performance:

```javascript
response.stats.totalTime; // seconds
response.stats.realTimeFactor; // synthesis time / audio duration; < 1 means streaming is possible
response.stats.audioDurationMs;
response.stats.totalSamples;
response.stats.tokensPerSecond;
response.stats.backendDevice; // 0 CPU, 1 GPU
response.stats.backendId; // 0 CPU, 1 Metal, 2 CUDA, 3 Vulkan, 4 OpenCL, 99 other
response.stats.enhancerBackendDevice; // -1 absent, 0 CPU, 1 GPU
response.stats.enhancerBackendId;
```

**Audio Format Specifications:**

* **Sample Rate:** engine-native by default (24 kHz for Chatterbox/CosyVoice3/MOSS,
  44.1 kHz for Supertonic/Parler/Audio8, 48 kHz with the enhancer), configurable
  via `config.outputSampleRate` (MOSS emits 24 kHz only)
* **Format:** 16-bit signed PCM, mono channel
* **Data Type:** Int16Array containing raw audio samples

## Public helpers and errors

Import `splitTtsText` from `@qvac/tts-ggml/text-chunker` and
`accumulateTextStream` from
`@qvac/tts-ggml/text-stream-accumulator`. The legacy
`@qvac/tts-ggml/lib/textStreamAccumulator.js` path remains supported.

`QvacErrorAddonTTSGgml` and `ERR_CODES` are exported from
`@qvac/tts-ggml`. Error codes are:

|  Code | Name                   |
| ----: | ---------------------- |
| 13001 | `FAILED_TO_ACTIVATE`   |
| 13002 | `FAILED_TO_APPEND`     |
| 13003 | `FAILED_TO_GET_STATUS` |
| 13004 | `FAILED_TO_PAUSE`      |
| 13005 | `FAILED_TO_CANCEL`     |
| 13006 | `FAILED_TO_DESTROY`    |
| 13007 | `FAILED_TO_UNLOAD`     |
| 13008 | `FAILED_TO_LOAD`       |
| 13009 | `FAILED_TO_RELOAD`     |
| 13010 | `FAILED_TO_STOP`       |
| 13011 | `JOB_ALREADY_RUNNING`  |

## More resources

[Package at npm](https://www.npmjs.com/package/@qvac/tts-ggml)
