---
title: "Text-to-Speech"
canonical: https://docs.qvac.tether.io/sdk/v0.20/ai-capabilities/text-to-speech/
collection: "SDK"
package: "@qvac/sdk"
line: v0.20
current_line: false
---

# Text-to-Speech (/sdk/v0.20/ai-capabilities/text-to-speech)



## Overview

Text-to-Speech uses [`@qvac/tts-ggml`](https://github.com/tetherto/qvac/tree/main/packages/tts-ggml) (GGML) as the inference engine. Load any supported model using `modelType: "tts"`. Then, provide `text` as input (with `inputType: "text"`) to generate speech audio.

`textToSpeech()` returns synchronously with `buffer`, `bufferStream`, `done`, `requestId`, `sampleRate`, `stats` and `stopReason`; see [Audio output](#audio-output), [Cancelling a run](#cancelling-a-run) and [Runtime statistics](#runtime-statistics) below.

## Functions

Use the following sequence of function calls:

1. [`loadModel()`](/sdk/v0.20/reference/api#loadmodel)
2. [`textToSpeech()`](/sdk/v0.20/reference/api#texttospeech)
3. [`unloadModel()`](/sdk/v0.20/reference/api#unloadmodel)

For how to use each function, see [SDK — API reference](/sdk/v0.20/reference/api/).

## Audio output

The SDK returns raw, mono, signed 16-bit PCM samples as a plain `number[]`. With `stream: false`, the `buffer` promise resolves to the complete sample array and `bufferStream` is empty. With `stream: true`, `buffer` resolves to `[]` and `bufferStream` yields individual samples as they become available. The data has no WAV header or other container metadata; use the [example utility](#utils) to write it as a WAV file.

`textToSpeech()` also returns a `sampleRate` promise, resolved from the first
audio frame the engine emits. Prefer it over inferring the rate from the load
config — `outputSampleRate` and the LavaSR enhancer both change it. It resolves
to `undefined` when the run produced no audio.

```ts
const result = textToSpeech({ modelId, text: "Hello.", stream: false });
const pcm = await result.buffer;
const sampleRate = await result.sampleRate; // e.g. 48000 with the enhancer
```

For reference, the rate each configuration produces:

| Configuration           | Output sample rate                                                                                                                                       |
| ----------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Chatterbox              | 24,000 Hz, or `modelConfig.outputSampleRate` when set                                                                                                    |
| Supertonic              | 44,100 Hz, or `modelConfig.outputSampleRate` when set                                                                                                    |
| Parler                  | 44,100 Hz, or `modelConfig.outputSampleRate` when native streaming is disabled; native streaming requires 44,100 Hz unless the LavaSR enhancer is loaded |
| CosyVoice3              | 24,000 Hz, or `modelConfig.outputSampleRate` when native streaming is disabled; native streaming requires 24,000 Hz unless the LavaSR enhancer is loaded |
| Audio8                  | 44,100 Hz, or `modelConfig.outputSampleRate` when set                                                                                                    |
| MOSS                    | 24,000 Hz only; MOSS takes no `outputSampleRate`                                                                                                         |
| LavaSR enhancer enabled | 48,000 Hz by default, or `modelConfig.outputSampleRate` when set                                                                                         |

## Cancelling a run

`textToSpeech()` surfaces a `requestId` synchronously on its result, and
`textToSpeechStream()` on its session. Pass it to
[`cancel()`](/sdk/v0.20/reference/api#cancel) to stop synthesis:

```ts
const result = textToSpeech({ modelId, text: longArticle });
// …later, e.g. the user navigated away
await cancel({ requestId: result.requestId });
```

<Callout type="warn">
  Abandoning `bufferStream` (or calling `destroy()` on a duplex session) stops
  delivery but leaves the native job running to completion. Only `cancel()`
  stops the engine.
</Callout>

`cancel({ modelId, kind: "tts" })` aborts the run in flight on that model
and every run queued behind it. To stop a single run, cancel by `requestId`:
synthesis is serialised per model (one active run, up to 256 waiting — beyond
that the call rejects with `RequestRejectedByPolicyError`), so cancelling a
queued run by id removes it from the queue without touching the one in flight.

A cancelled run ends cleanly rather than as an error: `done` resolves `true`,
`buffer` / `bufferStream` stop at whatever was delivered, and
`await result.stopReason` is `"cancelled"` (`"completed"` otherwise).

## Runtime statistics

`textToSpeech()` returns a `stats` promise carrying the engine's own runtime
statistics from the final frame (`undefined` if the run reported none).

<Callout type="info">
  The five backend fields (`backendDevice`, `backendId`, `gpuUnsupported`,
  `enhancerBackendDevice`, `enhancerBackendId`) are reported by batch
  synthesis (`stream: false`) only. The chunked paths — `stream: true` (the
  default), `sentenceStream: true`, and `textToSpeechStream()` — run one
  native job per sentence and aggregate only the timing and sample counts,
  plus Audio8's two Core ML fields, which keep the last sentence's value.
  MOSS is the exception for `stream: true`: it runs the whole text as one
  native job, so it reports them there too.
</Callout>

| Field                   | Meaning                                                                                                                      |
| ----------------------- | ---------------------------------------------------------------------------------------------------------------------------- |
| `audioDuration`         | Length of the synthesized audio, in milliseconds.                                                                            |
| `totalTime`             | Wall-clock synthesis time, in **seconds** (`audioDuration` is milliseconds).                                                 |
| `realTimeFactor`        | `totalTime × 1000 / audioDuration`; below 1.0 is faster than real time.                                                      |
| `tokensPerSecond`       | Input characters synthesized per second; on Audio8 and MOSS, codec frames per second.                                        |
| `totalSamples`          | PCM sample count.                                                                                                            |
| `generatedFrames`       | Audio8 and MOSS only: codec frames generated (Audio8 on a fixed 46 ms grid, MOSS at 12.5 per second).                        |
| `backendDevice`         | Compute device chosen at load: `0` CPU, `1` GPU.                                                                             |
| `backendId`             | Backend family: `0` CPU, `1` Metal, `2` CUDA, `3` Vulkan, `4` OpenCL, `99` other GPU.                                        |
| `gpuUnsupported`        | `1` when a GPU is present but unusable by engine policy, so the run fell back to CPU.                                        |
| `enhancerBackendDevice` | LavaSR enhancer device: `-1` not loaded, `0` CPU, `1` GPU.                                                                   |
| `enhancerBackendId`     | LavaSR enhancer backend family, same codes as `backendId`.                                                                   |
| `codecSidecarLoaded`    | Audio8 on macOS/iOS: `1` while the Core ML codec sidecar is attached, `0` without one or once a failing sidecar was retired. |
| `codecOnCoreml`         | Audio8: `1` when the synthesis ran its codec on the Core ML sidecar, `0` when it ran on the ggml backend `backendId` names.  |

```ts
const stats = await result.stats;
console.log(stats?.realTimeFactor, stats?.backendId, stats?.gpuUnsupported);
```

## Models

### Chatterbox

Chatterbox uses a **T3 GGUF** as the top-level `modelSrc` and an **S3Gen companion GGUF** via `modelConfig.s3genModelSrc`. Optional `referenceAudioSrc` supplies a WAV for voice cloning.

```ts
await loadModel({
  modelSrc: TTS_T3_TURBO_EN_CHATTERBOX_Q8_0,
  modelType: "tts",
  modelConfig: {
    ttsEngine: "chatterbox",
    language: "en",
    s3genModelSrc: TTS_S3GEN_EN_CHATTERBOX,
  },
});
```

<Callout type="info">
  Omitting 

  `ttsEngine`

   defaults to Chatterbox.
</Callout>

The multilingual (MTL) GGUFs need a tokenizer asset for two languages:
`mecabDictSrc` is required for Japanese (`language: "ja"`) and `cangjieTsvSrc`
for Chinese (`language: "zh"`). Both are rejected at load if missing.

`ttsSpeed` is Chatterbox's speech-rate control — a pitch-preserving WSOLA
time-stretch bounded to `0.25`–`4.0` (Chatterbox has no `pace` channel).
`nCtx` caps the T3 context length in tokens (\~25 tokens ≈ 1 s of audio); the
KV cache is allocated up front at that length, so it directly bounds memory.
`kvCacheType` picks the KV-cache dtype — `f16` (the safe cross-backend
default), `f32`, or `q8_0` (smaller and faster where the backend implements
its ops). `seed`, `threads`, `nGpuLayers`, `outputSampleRate` and `backendsDir` work
as on the other engines; `openclCacheDir` (the Android OpenCL program cache)
is accepted by Chatterbox, Supertonic and CosyVoice3 only. Chatterbox supports
[LavaSR](#lavasr) post-processing.

### Supertonic

Supertonic uses a **single GGUF** via top-level `modelSrc`. Set `voice`, `ttsSpeed`, and `ttsNumInferenceSteps` in `modelConfig` as needed. Multilingual output is selected by the GGUF (e.g. `TTS_MULTILINGUAL_SUPERTONIC2_Q8_0`) plus `language`.

```ts
await loadModel({
  modelSrc: TTS_EN_SUPERTONIC_Q8_0,
  modelType: "tts",
  modelConfig: {
    ttsEngine: "supertonic",
    language: "en",
    voice: "F1",
  },
});
```

Supertonic takes its speech rate either as an exact multiplier (`ttsSpeed`) or
as the canonical `pace` vocabulary (`slow` / `moderate` / `fast`), which maps
onto the duration multiplier relative to the GGUF's own default. The two are
mutually exclusive and rejected together at load. Unlike Parler and CosyVoice3,
Supertonic conditions `pace` when the engine is built, so it cannot be changed
per request — set it in `modelConfig`.

`seed`, `threads`, `nGpuLayers`, `outputSampleRate` and `backendsDir` work
as on the other engines; `openclCacheDir` is accepted by Chatterbox,
Supertonic and CosyVoice3 only, and `vulkanCacheDir` additionally persists
the compiled Vulkan pipeline cache when `useGPU` is set.

### Parler-TTS

Parler-TTS uses a **single GGUF** and conditions speech from either a free-text
description or structured voice fields. Load-time fields provide defaults;
`textToSpeech()` and `textToSpeechStream()` can override them per request.
The registry includes Mini v1, Large v1, and Indic variants; the Indic model
supports 21 languages and script-native digit normalization.

```ts
const modelId = await loadModel({
  modelSrc: TTS_MINI_V1_EN_PARLER_TTS_Q8_0,
  modelType: "tts",
  modelConfig: {
    ttsEngine: "parler",
    voice: "Laura",
    seed: 42,
    topK: 1,
  },
});

const result = textToSpeech({
  modelId,
  text: "Welcome to Parler text-to-speech.",
  inputType: "text",
  stream: false,
  emotion: "happy",
  pace: "moderate",
});

const pcm = await result.buffer;
```

For free-form conditioning, use `description` or its alias
`voiceDescription`. Otherwise, compose a description from `voice`, `emotion`,
`pitch`, `pace`, `expressivity`, `noise`, `reverb`, and `quality`.

<Callout type="warn">
  `description` and `voiceDescription` cannot be combined with each other or
  with the structured voice fields. `description`, `voiceDescription` and the
  descriptor fields (`voice`, `pitch`, `expressivity`, `noise`, `reverb`,
  `quality`) are Parler-only; `emotion` and `pace` are also accepted per
  request by CosyVoice3 (emotion limited to `anger`, `happy`, `neutral`,
  `sad`). Sending any of them to Chatterbox, Supertonic or Audio8 returns a
  request-validation error.
</Callout>

Supported emotions are `command`, `anger`, `narration`, `conversation`,
`disgust`, `fear`, `happy`, `neutral`, `proper noun`, `news`, `sad`, and
`surprise`.

Parler supports all three SDK streaming surfaces:

* `textToSpeech({ stream: true })` for incremental PCM samples.
* `textToSpeech({ stream: true, sentenceStream: true })` for PCM plus
  sentence/chunk metadata.
* `textToSpeechStream()` when text itself arrives incrementally.

Set `streamChunkTokens` above zero to enable native chunk streaming;
`streamFirstChunkTokens` only tunes the first chunk and does not enable
streaming by itself. Native streaming emits at 44.1 kHz, so omit
`outputSampleRate`, set it to `44100`, or load the LavaSR enhancer (which
resamples seam-free) when `streamChunkTokens > 0`.
Integer generation controls use signed 32-bit values. Parler also supports
[LavaSR](#lavasr) post-processing via `lavasrEnhancerModelSrc` and
`lavasrDenoiserModelSrc` (the denoiser is batch-only — disable native chunk
streaming to use it).

The generated Python client uses the same contract:

```python
from tetherto.qvac_sdk import TextToSpeechRequest, load_model, text_to_speech
from tetherto.qvac_sdk.models import TTS_MINI_V1_EN_PARLER_TTS_Q8_0

model_id = await load_model(
    transport,
    model_src=TTS_MINI_V1_EN_PARLER_TTS_Q8_0,
    model_config={
        "ttsEngine": "parler",
        "voice": "Laura",
        "seed": 42,
        "topK": 1,
    },
)

request = TextToSpeechRequest.model_validate({
    "type": "textToSpeech",
    "modelId": model_id,
    "text": "Welcome to Parler text-to-speech.",
    "inputType": "text",
    "stream": False,
    "emotion": "happy",
})

samples = []
async for response in text_to_speech(transport, request):
    samples.extend(response.buffer)
```

### CosyVoice3

CosyVoice3 uses the **LLM GGUF** as the top-level `modelSrc`; its registry
companion set auto-downloads the flow and HiFT GGUFs, the baked `voice.gguf`,
and the tokenizer files alongside it. Speech is synthesized with the baked
default voice unless you supply a reference recording (see
[voice cloning](#cosyvoice3-voice-cloning) below), and emits native 24 kHz
audio.

```ts
const modelId = await loadModel({
  modelSrc: TTS_COSYVOICE3_LLM_COSYVOICE_Q8_0,
  modelType: "tts",
  modelConfig: {
    ttsEngine: "cosyvoice3",
    seed: 42,
  },
});

const result = textToSpeech({
  modelId,
  text: "Hey, how are you doing today?",
  inputType: "text",
  stream: false,
  emotion: "happy",
});

const pcm = await result.buffer;
```

Conditioning uses one of three controls:

* `emotion` — `anger`, `happy`, `neutral`, or `sad`.
* `pace` — `slow`, `moderate`, or `fast`; `moderate` engages nothing and
  takes the plain zero-shot path.
* `instruct` — a raw instruction string, or a structured object with
  `dialect`, `volume` (`loud` / `soft`), or `style` (`peppa` / `robot`),
  resolved by precedence `dialect > volume > style`. Supported dialects:
  `cantonese`, `northeastern`, `gansu`, `guizhou`, `henan`, `hubei`, `hunan`,
  `jiangxi`, `minnan`, `ningxia`, `shanxi`, `shaanxi`, `shandong`,
  `shanghai`, `sichuan`, `tianjin`, and `yunnan`.

<Callout type="warn">
  CosyVoice3 accepts **one conditioning control per synthesis** — an
  `emotion`, a non-`moderate` `pace`, or an `instruct`. Combining them is
  rejected at load and request validation. `emotion` and `pace` can also be
  set per request on `textToSpeech()`; `instruct` is load-time only.
</Callout>

Set `streamChunkTokens` above zero for native chunk streaming
(`streamFirstChunkTokens` tunes the first chunk only). Native streaming emits
at 24 kHz, so omit `outputSampleRate` or set it to `24000` unless the LavaSR
enhancer is loaded. `useGPU` offloads to Metal on macOS/iOS, Vulkan on
desktop Linux/Windows, and OpenCL on Adreno Android devices; `threads`,
`nGpuLayers`, `seed`, and `outputSampleRate` (8,000–192,000 Hz) work as on
the other engines. CosyVoice3 also supports [LavaSR](#lavasr)
post-processing via `lavasrEnhancerModelSrc` (48 kHz bandwidth extension) and
`lavasrDenoiserModelSrc` (batch-only; runs before the enhancer).

#### CosyVoice3 voice cloning

Supply `referenceAudioSrc` — a WAV of 0.5–30 s (5–15 s of clean speech is the
sweet spot; multichannel input is downmixed to mono) — to replace the baked
voice with the speaker in the recording. Cloning needs two extra GGUFs that are
**not** part of the LLM's companion set, because they are a \~300 MB opt-in:

* `cosyvoice3S3tokModelSrc` — the speech tokenizer (`speech_tokenizer_v3`).
  Constants: `TTS_COSYVOICE3_S3TOK_COSYVOICE_Q8_0`,
  `TTS_COSYVOICE3_S3TOK_COSYVOICE_FP16`, `TTS_COSYVOICE3_S3TOK_COSYVOICE_FP32`.
* `cosyvoice3CampplusModelSrc` — the CAM++ speaker encoder. Constant:
  `TTS_COSYVOICE3_CAMPPLUS_COSYVOICE_FP32`.

`promptText` selects the cloning mode: set it to the verbatim transcript of the
recording for **zero-shot** cloning (best fidelity in the reference's own
language), or omit it for **cross-lingual** cloning (timbre only, so you can
synthesize a different language than the reference speaks). Cloning composes
with `instruct`.

```ts
const modelId = await loadModel({
  modelSrc: TTS_COSYVOICE3_LLM_COSYVOICE_Q8_0,
  modelType: "tts",
  modelConfig: {
    ttsEngine: "cosyvoice3",
    referenceAudioSrc: "/path/to/reference.wav",
    cosyvoice3S3tokModelSrc: TTS_COSYVOICE3_S3TOK_COSYVOICE_Q8_0.src,
    cosyvoice3CampplusModelSrc: TTS_COSYVOICE3_CAMPPLUS_COSYVOICE_FP32.src,
    // Zero-shot: the exact words spoken in reference.wav.
    // Omit promptText for cross-lingual cloning.
    promptText: "This is exactly what the reference recording says.",
  },
});
```

<Callout type="warn">
  The reference is baked at load. All three fields are validated together — a
  reference without either cloning GGUF is rejected at load rather than
  silently falling back to the baked voice. Changing voices means loading the
  model again.
</Callout>

### Audio8

Audio8 uses the **DualAR LM GGUF** as the top-level `modelSrc`
(`TTS_LM_MULTILINGUAL_AUDIO8_Q8_0` or `TTS_LM_MULTILINGUAL_AUDIO8_FP16`) and
a required **codec decoder** via `modelConfig.audio8CodecDecoderModelSrc`.
The model is multilingual — the language is inferred from the text itself —
and emits native 44.1 kHz audio.

```ts
const modelId = await loadModel({
  modelSrc: TTS_LM_MULTILINGUAL_AUDIO8_Q8_0,
  modelType: "tts",
  modelConfig: {
    ttsEngine: "audio8",
    audio8CodecDecoderModelSrc: TTS_CODEC_DECODER_AUDIO8_Q8_0,
    temperature: 0.7,
    topP: 0.9,
    seed: 42,
  },
});

const result = textToSpeech({
  modelId,
  text: "Hey, how are you doing today?",
  inputType: "text",
  stream: false,
});

const pcm = await result.buffer;
```

For zero-shot voice cloning, add the optional **codec encoder**
(`audio8CodecEncoderModelSrc`, e.g. `TTS_CODEC_ENCODER_AUDIO8_Q8_0`) and
supply `referenceAudioSrc` (a reference recording) together with
`referenceText` (its exact transcript). The three fields are validated
together — a reference without its transcript or without the encoder is
rejected at load.

Generation is sampled by default: `temperature` (default `0.7`), `topK`
(`50`), and `topP` (`0.9`) tune it, `greedy: true` takes the argmax instead
and ignores all three, and `maxFrames` caps generation in codec frames
(\~21.5 frames per second of audio). `seed`, `threads`, `nGpuLayers`, and
`outputSampleRate` work as on the other engines; `useGPU` offloads to Metal
on macOS/iOS, Vulkan on desktop Linux/Windows, and OpenCL on Android/Adreno —
elsewhere it falls back to CPU and sets `stats.gpuUnsupported` in the
response.

On macOS and iOS, loading the codec decoder from its registry constant also
downloads its Apple Core ML bundle (`audio8-codec-decoder.mlmodelc`, about
150 MB, the same bundle for every decoder tier) and places it beside the
decoder GGUF.
The codec's synthesis stack — the upsampling stages and the DAC decoder — then
runs on Core ML, while the language model stays on `useGPU`'s backend. Other
platforms download only the GGUF. If the bundle is unavailable at load, or a
synthesis fails on it, Audio8 runs the codec on ggml instead; the
`codecSidecarLoaded` and `codecOnCoreml` stats report which path ran.

<Callout type="warn">
  Audio8 has no `emotion`, `pace`, or description fields (at load or per
  request), no LavaSR post-processing, and no native chunk streaming.
  Sentence streaming (`sentenceStream: true`) and duplex
  `textToSpeechStream()` work as with every engine.
</Callout>

For model constants, see [SDK — Models](/sdk/v0.20/#models).

### MOSS

MOSS runs OpenMOSS **MOSS-TTS v1.5** (single speaker) and **MOSS-TTSD**
(multi-speaker dialogue) on the MOSS Delay engine. The top-level `modelSrc` is
the **backbone GGUF** (`TTS_DELAY_LLM_MOSS_TTS_F16`, or a local
`moss-ttsd-*.gguf` for dialogue) and a **codec decoder** is required via
`modelConfig.mossCodecDecoderModelSrc` (`TTS_CODEC_DECODER_MOSS_TTS_F16`).
Output is native 24 kHz. The backbones have 8B parameters (about 17 GB at
f16), so MOSS targets desktop hosts; mobile is not supported.

```ts
const modelId = await loadModel({
  modelSrc: TTS_DELAY_LLM_MOSS_TTS_F16,
  modelType: "tts",
  modelConfig: {
    ttsEngine: "moss",
    language: "en",
    mossCodecDecoderModelSrc: TTS_CODEC_DECODER_MOSS_TTS_F16,
    streamChunkTokens: 25,
    useGPU: true,
  },
});

const result = textToSpeech({
  modelId,
  text: "Hold on [pause 1.0s] here it comes.",
});

for await (const sample of result.bufferStream) {
  // Chunks of 25 codec frames (2 s) arrive while the backbone is still generating.
}
```

The speech is directable from the text itself: `[pause 2.0s]` markers insert
a silence of roughly that length, and inline Pinyin (`ni3 hao3`) or IPA
(`/həloʊ/`) steers pronunciation. `durationTokens` asks for a target length in
codec frames (12.5 per second, so `38` is about 3 s; `0` or unset keeps the
length free; at most `2015`). `language` is written into the model prompt as a
hint (default `en`).

MOSS streams natively. With `streamChunkTokens` set, `stream: true` (the
default) delivers audio every that-many codec frames while generation
continues; without it the audio arrives in one piece when synthesis ends.
Either way the text is synthesized as one utterance rather than sentence by
sentence, so `durationTokens` and the pause markers apply to all of it.
`sentenceStream: true` and `textToSpeechStream()` still split a single-speaker
text into sentences; `sentenceStreamLocale` and `sentenceStreamMaxChunkScalars`
only apply there, and a MOSS request without `sentenceStream: true` rejects
them.

`seed` (engine default `1234`), `threads` and `backendsDir` work as on the
other engines. MOSS runs on the CPU by default; `useGPU: true` or a non-zero
`nGpuLayers` asks for a GPU backend, and when none is usable it falls back to
the CPU and sets `stats.gpuUnsupported`.

#### MOSS voice cloning

Add the **codec encoder** (`mossCodecEncoderModelSrc`,
`TTS_CODEC_ENCODER_MOSS_TTS_F16`) and a reference recording
(`referenceAudioSrc`). Unlike Audio8, MOSS needs no transcript. The
recording must already be sampled at **24 kHz** (there is no resampling;
multichannel audio is downmixed, and recordings are capped at 60 s). The voice is
encoded once at load and fixed for the loaded model, so changing voices means
loading again.

```ts
const modelId = await loadModel({
  modelSrc: TTS_DELAY_LLM_MOSS_TTS_F16,
  modelType: "tts",
  modelConfig: {
    ttsEngine: "moss",
    mossCodecDecoderModelSrc: TTS_CODEC_DECODER_MOSS_TTS_F16,
    mossCodecEncoderModelSrc: TTS_CODEC_ENCODER_MOSS_TTS_F16,
    referenceAudioSrc: "/voices/speaker-24k.wav",
  },
});
```

#### MOSS-TTSD dialogue

The MOSS-TTSD backbone is not in the QVAC Registry yet, so pass it as a local
file; it shares the codec halves with MOSS-TTS. Load it with
`dialogueReferenceSrcs`: one 24 kHz recording per speaker, in the order the
text tags them (`[S1]`, `[S2]`, …; one to five speakers, at most 60 s of
reference audio combined). The model continues the references, so the text
must open with what each recording says under its tag, followed by the lines
to generate; only the new lines come out as audio. Dialogue needs the codec
encoder and excludes `referenceAudioSrc`.

```ts
const modelId = await loadModel({
  modelSrc: "/models/moss-ttsd-f16.gguf",
  modelType: "tts",
  modelConfig: {
    ttsEngine: "moss",
    mossCodecDecoderModelSrc: TTS_CODEC_DECODER_MOSS_TTS_F16,
    mossCodecEncoderModelSrc: TTS_CODEC_ENCODER_MOSS_TTS_F16,
    dialogueReferenceSrcs: ["/voices/alice.wav", "/voices/bob.wav"],
    streamChunkTokens: 25,
  },
});

const result = textToSpeech({
  modelId,
  text:
    "[S1] What alice.wav says. [S2] What bob.wav says. " +
    "[S1] Did the build finish? [S2] Yes, every test passed.",
});
```

<Callout type="warn">
  A dialogue cannot be split into sentences, because every synthesis has to
  open with the reference transcripts: `sentenceStream: true` and
  `textToSpeechStream()` are rejected on a dialogue model. Use
  `textToSpeech()` with `stream: true` and `streamChunkTokens` for chunked
  audio. MOSS has no `emotion`, `pace`, or description fields, no LavaSR
  post-processing, and no `streamFirstChunkTokens`.
</Callout>

### LavaSR

With Chatterbox, Supertonic, Parler, or CosyVoice3, you can choose to perform post-processing using **LavaSR** models. They are applied to the synthesized audio before it is returned, and each stage is enabled purely by supplying its model source — i.e., there is no separate on/off flag. (Audio8 and MOSS do not support them.)

* `lavasrDenoiserModelSrc` — a denoiser GGUF that cleans the speech (noise reduction). It runs **first** and is rate-preserving. Constants: `TTS_DENOISER_LAVASR_FP16`, `TTS_DENOISER_LAVASR_FP32`.
* `lavasrEnhancerModelSrc` — an enhancer GGUF that neurally bandwidth-extends the output to **48 kHz**. It runs **after** the denoiser. Constants: `TTS_ENHANCER_LAVASR_FP16`, `TTS_ENHANCER_LAVASR_FP32`.

<Callout type="warn">
  The denoiser needs the whole utterance, so it cannot be combined with native
  chunk streaming (`streamChunkTokens > 0`) on Chatterbox, Parler, or
  CosyVoice3. The combination is rejected at load. The enhancer has no such
  restriction — it resamples seam-free across chunk boundaries.
</Callout>

```ts
await loadModel({
  modelSrc: TTS_MULTILINGUAL_SUPERTONIC3_Q8_0,
  modelType: "tts",
  modelConfig: {
    ttsEngine: "supertonic",
    language: "en",
    voice: "F1",
    // Denoiser runs first (rate-preserving)…
    lavasrDenoiserModelSrc: TTS_DENOISER_LAVASR_FP16.src,
    // …then the enhancer bandwidth-extends to 48 kHz.
    lavasrEnhancerModelSrc: TTS_ENHANCER_LAVASR_FP16.src,
  },
});
```

## Examples

### Chatterbox

The following script shows an example of Chatterbox TTS with voice cloning from a reference audio file. Use it with [`utils.js` / `utils.ts`](#utils):

<Tabs>
  <Tab value="js" label="JavaScript" default>
    <WrapCode>
      ```js file=<rootDir>/packages/sdk/dist/examples/tts/chatterbox.js title="tts-chatterbox.js" lineNumbers
      import { loadModel, textToSpeech, unloadModel, TTS_T3_TURBO_EN_CHATTERBOX_Q8_0, TTS_S3GEN_EN_CHATTERBOX } from '@qvac/sdk';
      import { createWav, playAudio, int16ArrayToBuffer, createWavHeader } from './utils';
      // Chatterbox TTS (GGML): voice cloning with optional reference audio.
      // Uses registry model constants — downloads automatically from QVAC Registry.
      // Usage: node chatterbox.ts [referenceAudioSrc]
      const [referenceAudioSrc] = process.argv.slice(2);
      // Only a fallback: the engine reports the rate it actually produced, and
      // `outputSampleRate` (plus the LavaSR enhancer) can move it off this default.
      const CHATTERBOX_SAMPLE_RATE = 24000;
      try {
          const modelId = await loadModel({
              modelSrc: TTS_T3_TURBO_EN_CHATTERBOX_Q8_0,
              modelConfig: {
                  ttsEngine: 'chatterbox',
                  language: 'en',
                  s3genModelSrc: TTS_S3GEN_EN_CHATTERBOX.src,
                  streamChunkTokens: 25,
                  streamFirstChunkTokens: 10,
                  cfmSteps: 1,
                  threads: 8,
                  ...(referenceAudioSrc ? { referenceAudioSrc } : {})
              },
              onProgress: (p) => {
                  const mb = (n) => (n / 1e6).toFixed(1);
                  const line = `▸ Downloading ${p.percentage.toFixed(0)}% (${mb(p.downloaded)}/${mb(p.total)} MB)`;
                  process.stderr.write(process.stderr.isTTY ? `\r${line}` : `${line}\n`);
                  if (p.percentage >= 100)
                      process.stderr.write('\n');
              }
          });
          console.log(`▸ Model loaded: ${modelId}`);
          console.log('▸ Testing Text-to-Speech...');
          const result = textToSpeech({
              modelId,
              text: `QVAC SDK is the canonical entry point to QVAC. Written in TypeScript, it provides all QVAC capabilities through a unified interface while also abstracting away the complexity of running your application in a JS environment other than Bare. Supported JS environments include Bare, Node.js, Expo and Bun.`,
              inputType: 'text',
              stream: false
          });
          const audioBuffer = await result.buffer;
          const sampleRate = (await result.sampleRate) ?? CHATTERBOX_SAMPLE_RATE;
          console.log(`▸ TTS complete. Total bytes: ${audioBuffer.length}`);
          console.log('▸ Saving audio to file...');
          createWav(audioBuffer, sampleRate, 'tts-output.wav');
          console.log('▸ Audio saved to tts-output.wav');
          console.log('▸ Playing audio...');
          const audioData = int16ArrayToBuffer(audioBuffer);
          const wavBuffer = Buffer.concat([createWavHeader(audioData.length, sampleRate), audioData]);
          playAudio(wavBuffer);
          console.log('▸ Audio playback complete');
          await unloadModel({ modelId });
          console.log('▸ Model unloaded');
          process.exit(0);
      }
      catch (error) {
          console.error('✖', error);
          process.exit(1);
      }
      ```
    </WrapCode>
  </Tab>

  <Tab value="ts" label="TypeScript">
    <WrapCode>
      ```ts file=<rootDir>/packages/sdk/examples/tts/chatterbox.ts title="tts-chatterbox.ts" lineNumbers
      import {
        loadModel,
        textToSpeech,
        unloadModel,
        type ModelProgressUpdate,
        TTS_T3_TURBO_EN_CHATTERBOX_Q8_0,
        TTS_S3GEN_EN_CHATTERBOX
      } from '@qvac/sdk'
      import { createWav, playAudio, int16ArrayToBuffer, createWavHeader } from './utils'

      // Chatterbox TTS (GGML): voice cloning with optional reference audio.
      // Uses registry model constants — downloads automatically from QVAC Registry.
      // Usage: node chatterbox.ts [referenceAudioSrc]
      const [referenceAudioSrc] = process.argv.slice(2)

      // Only a fallback: the engine reports the rate it actually produced, and
      // `outputSampleRate` (plus the LavaSR enhancer) can move it off this default.
      const CHATTERBOX_SAMPLE_RATE = 24000

      try {
        const modelId = await loadModel({
          modelSrc: TTS_T3_TURBO_EN_CHATTERBOX_Q8_0,
          modelConfig: {
            ttsEngine: 'chatterbox',
            language: 'en',
            s3genModelSrc: TTS_S3GEN_EN_CHATTERBOX.src,
            streamChunkTokens: 25,
            streamFirstChunkTokens: 10,
            cfmSteps: 1,
            threads: 8,
            ...(referenceAudioSrc ? { referenceAudioSrc } : {})
          },
          onProgress: (p: ModelProgressUpdate) => {
            const mb = (n: number) => (n / 1e6).toFixed(1)
            const line = `▸ Downloading ${p.percentage.toFixed(0)}% (${mb(p.downloaded)}/${mb(p.total)} MB)`
            process.stderr.write(process.stderr.isTTY ? `\r${line}` : `${line}\n`)
            if (p.percentage >= 100) process.stderr.write('\n')
          }
        })

        console.log(`▸ Model loaded: ${modelId}`)

        console.log('▸ Testing Text-to-Speech...')
        const result = textToSpeech({
          modelId,
          text: `QVAC SDK is the canonical entry point to QVAC. Written in TypeScript, it provides all QVAC capabilities through a unified interface while also abstracting away the complexity of running your application in a JS environment other than Bare. Supported JS environments include Bare, Node.js, Expo and Bun.`,
          inputType: 'text',
          stream: false
        })

        const audioBuffer = await result.buffer
        const sampleRate = (await result.sampleRate) ?? CHATTERBOX_SAMPLE_RATE
        console.log(`▸ TTS complete. Total bytes: ${audioBuffer.length}`)

        console.log('▸ Saving audio to file...')
        createWav(audioBuffer, sampleRate, 'tts-output.wav')
        console.log('▸ Audio saved to tts-output.wav')

        console.log('▸ Playing audio...')
        const audioData = int16ArrayToBuffer(audioBuffer)
        const wavBuffer = Buffer.concat([createWavHeader(audioData.length, sampleRate), audioData])
        playAudio(wavBuffer)
        console.log('▸ Audio playback complete')

        await unloadModel({ modelId })
        console.log('▸ Model unloaded')
        process.exit(0)
      } catch (error) {
        console.error('✖', error)
        process.exit(1)
      }
      ```
    </WrapCode>
  </Tab>
</Tabs>

### Supertonic

The following script shows an example of Supertonic TTS for general-purpose speech synthesis. Use it with [`utils.js` / `utils.ts`](#utils):

<Tabs>
  <Tab value="js" label="JavaScript" default>
    <WrapCode>
      ```js file=<rootDir>/packages/sdk/dist/examples/tts/supertonic.js title="tts-supertonic.js" lineNumbers
      import { loadModel, textToSpeech, unloadModel, TTS_MULTILINGUAL_SUPERTONIC3_Q8_0 } from '@qvac/sdk';
      import { createWav, playAudio, int16ArrayToBuffer, createWavHeader } from './utils';
      // Supertonic 3 TTS (GGML): fast multilingual synthesis with baked-in voices.
      // Uses registry model constants — downloads automatically from QVAC Registry.
      // Only a fallback: the engine reports the rate it actually produced, and
      // `outputSampleRate` (plus the LavaSR enhancer) can move it off this default.
      const SUPERTONIC_SAMPLE_RATE = 44100;
      try {
          const modelId = await loadModel({
              modelSrc: TTS_MULTILINGUAL_SUPERTONIC3_Q8_0,
              modelConfig: {
                  ttsEngine: 'supertonic',
                  language: 'en',
                  voice: 'F1',
                  ttsSpeed: 1.05,
                  ttsNumInferenceSteps: 5
              },
              onProgress: (p) => {
                  const mb = (n) => (n / 1e6).toFixed(1);
                  const line = `▸ Downloading ${p.percentage.toFixed(0)}% (${mb(p.downloaded)}/${mb(p.total)} MB)`;
                  process.stderr.write(process.stderr.isTTY ? `\r${line}` : `${line}\n`);
                  if (p.percentage >= 100)
                      process.stderr.write('\n');
              }
          });
          console.log(`▸ Model loaded: ${modelId}`);
          console.log('▸ Testing Text-to-Speech...');
          const result = textToSpeech({
              modelId,
              text: `QVAC SDK is the canonical entry point to QVAC. Written in TypeScript, it provides all QVAC capabilities through a unified interface while also abstracting away the complexity of running your application in a JS environment other than Bare. Supported JS environments include Bare, Node.js, Expo and Bun.`,
              inputType: 'text',
              stream: false
          });
          const audioBuffer = await result.buffer;
          const sampleRate = (await result.sampleRate) ?? SUPERTONIC_SAMPLE_RATE;
          console.log(`▸ TTS complete. Total samples: ${audioBuffer.length}`);
          console.log('▸ Saving audio to file...');
          createWav(audioBuffer, sampleRate, 'supertonic-output.wav');
          console.log('▸ Audio saved to supertonic-output.wav');
          console.log('▸ Playing audio...');
          const audioData = int16ArrayToBuffer(audioBuffer);
          const wavBuffer = Buffer.concat([createWavHeader(audioData.length, sampleRate), audioData]);
          playAudio(wavBuffer);
          console.log('▸ Audio playback complete');
          await unloadModel({ modelId });
          console.log('▸ Model unloaded');
          process.exit(0);
      }
      catch (error) {
          console.error('✖', error);
          process.exit(1);
      }
      ```
    </WrapCode>
  </Tab>

  <Tab value="ts" label="TypeScript">
    <WrapCode>
      ```ts file=<rootDir>/packages/sdk/examples/tts/supertonic.ts title="tts-supertonic.ts" lineNumbers
      import {
        loadModel,
        textToSpeech,
        unloadModel,
        type ModelProgressUpdate,
        TTS_MULTILINGUAL_SUPERTONIC3_Q8_0
      } from '@qvac/sdk'
      import { createWav, playAudio, int16ArrayToBuffer, createWavHeader } from './utils'

      // Supertonic 3 TTS (GGML): fast multilingual synthesis with baked-in voices.
      // Uses registry model constants — downloads automatically from QVAC Registry.
      // Only a fallback: the engine reports the rate it actually produced, and
      // `outputSampleRate` (plus the LavaSR enhancer) can move it off this default.
      const SUPERTONIC_SAMPLE_RATE = 44100

      try {
        const modelId = await loadModel({
          modelSrc: TTS_MULTILINGUAL_SUPERTONIC3_Q8_0,
          modelConfig: {
            ttsEngine: 'supertonic',
            language: 'en',
            voice: 'F1',
            ttsSpeed: 1.05,
            ttsNumInferenceSteps: 5
          },
          onProgress: (p: ModelProgressUpdate) => {
            const mb = (n: number) => (n / 1e6).toFixed(1)
            const line = `▸ Downloading ${p.percentage.toFixed(0)}% (${mb(p.downloaded)}/${mb(p.total)} MB)`
            process.stderr.write(process.stderr.isTTY ? `\r${line}` : `${line}\n`)
            if (p.percentage >= 100) process.stderr.write('\n')
          }
        })

        console.log(`▸ Model loaded: ${modelId}`)

        console.log('▸ Testing Text-to-Speech...')
        const result = textToSpeech({
          modelId,
          text: `QVAC SDK is the canonical entry point to QVAC. Written in TypeScript, it provides all QVAC capabilities through a unified interface while also abstracting away the complexity of running your application in a JS environment other than Bare. Supported JS environments include Bare, Node.js, Expo and Bun.`,
          inputType: 'text',
          stream: false
        })

        const audioBuffer = await result.buffer
        const sampleRate = (await result.sampleRate) ?? SUPERTONIC_SAMPLE_RATE
        console.log(`▸ TTS complete. Total samples: ${audioBuffer.length}`)

        console.log('▸ Saving audio to file...')
        createWav(audioBuffer, sampleRate, 'supertonic-output.wav')
        console.log('▸ Audio saved to supertonic-output.wav')

        console.log('▸ Playing audio...')
        const audioData = int16ArrayToBuffer(audioBuffer)
        const wavBuffer = Buffer.concat([createWavHeader(audioData.length, sampleRate), audioData])
        playAudio(wavBuffer)
        console.log('▸ Audio playback complete')

        await unloadModel({ modelId })
        console.log('▸ Model unloaded')
        process.exit(0)
      } catch (error) {
        console.error('✖', error)
        process.exit(1)
      }
      ```
    </WrapCode>
  </Tab>

  <Tab value="python" label="Python">
    <WrapCode>
      ```python file=<rootDir>/packages/sdk-python/examples/text_to_speech.py title="text_to_speech.py" lineNumbers
      """Python port of packages/sdk/examples/tts/supertonic.ts.

      Text-to-speech with Supertonic 3 (GGML). `text_to_speech` is a server-stream;
      in non-stream mode it yields the full audio as one frame. Samples come back as
      int16 PCM values, written here to a WAV with the stdlib `wave` module (the JS
      example's `createWav`/`playAudio` helpers).

      RUN: python examples/text_to_speech.py
      """

      from __future__ import annotations

      import array
      import asyncio
      import sys
      import wave

      from tetherto.qvac_sdk import (
          Client,
          TextToSpeechRequest,
          load_model,
          text_to_speech,
          unload_model,
      )
      from tetherto.qvac_sdk.models import TTS_MULTILINGUAL_SUPERTONIC3_Q8_0

      SUPERTONIC_SAMPLE_RATE = 44100


      def print_progress(p) -> None:
          """Print model download progress; pass as `on_progress=` to `load_model`."""
          line = (
              f"▸ Downloading {p.percentage:.0f}% "
              f"({p.downloaded / 1e6:.1f}/{p.total / 1e6:.1f} MB)"
          )
          print(line, end="\r" if sys.stderr.isatty() else "\n", file=sys.stderr)
          if p.percentage >= 100:
              print(file=sys.stderr)


      def write_wav(samples, sample_rate, path) -> None:
          pcm = array.array("h", (max(-32768, min(32767, int(s))) for s in samples))
          with wave.open(path, "wb") as w:
              w.setnchannels(1)
              w.setsampwidth(2)
              w.setframerate(sample_rate)
              w.writeframes(pcm.tobytes())


      async def main() -> int:
          async with Client() as client:
              t = client.transport
              try:
                  model_id = await load_model(
                      t,
                      model_src=TTS_MULTILINGUAL_SUPERTONIC3_Q8_0,
                      model_config={
                          "ttsEngine": "supertonic",
                          "language": "en",
                          "voice": "F1",
                          "ttsSpeed": 1.05,
                          "ttsNumInferenceSteps": 5,
                      },
                      on_progress=print_progress,
                  )
                  print(f"▸ Model loaded: {model_id}")

                  print("▸ Testing Text-to-Speech...")
                  request = TextToSpeechRequest.model_validate(
                      {
                          "type": "textToSpeech",
                          "modelId": model_id,
                          "text": (
                              "QVAC SDK is the canonical entry point to QVAC. It provides all "
                              "QVAC capabilities through a unified interface."
                          ),
                          "inputType": "text",
                          "stream": False,
                      }
                  )

                  samples: list[float] = []
                  async for response in text_to_speech(t, request):
                      samples.extend(response.buffer)
                  print(f"▸ TTS complete. Total samples: {len(samples)}")

                  out = "supertonic-output.wav"
                  write_wav(samples, SUPERTONIC_SAMPLE_RATE, out)
                  print(f"▸ Audio saved to {out}")

                  await unload_model(t, model_id)
                  print("▸ Model unloaded")
              except Exception as error:
                  print(f"✖ {error}", file=sys.stderr)
                  return 1
          return 0


      if __name__ == "__main__":
          sys.exit(asyncio.run(main()))
      ```
    </WrapCode>
  </Tab>
</Tabs>

### Parler-TTS

The following TypeScript example loads the registry-hosted Parler Mini v1
model, applies per-request emotion conditioning, saves the PCM output, and
plays it:

<Tabs>
  <Tab value="js" label="JavaScript" default>
    <WrapCode>
      ```js file=<rootDir>/packages/sdk/dist/examples/tts/parler.js title="tts-parler.js" lineNumbers
      import { loadModel, textToSpeech, unloadModel, TTS_MINI_V1_EN_PARLER_TTS_Q8_0 } from '@qvac/sdk';
      import { createWav, playAudio, int16ArrayToBuffer, createWavHeader } from './utils';
      // Parler-TTS (GGML): description-conditioned speech with per-call voice controls.
      // Uses the registry-hosted Mini v1 Q8_0 model and its native 44.1 kHz output.
      // Only a fallback: the engine reports the rate it actually produced, and
      // `outputSampleRate` (plus the LavaSR enhancer) can move it off this default.
      const PARLER_SAMPLE_RATE = 44100;
      try {
          const modelId = await loadModel({
              modelSrc: TTS_MINI_V1_EN_PARLER_TTS_Q8_0,
              modelConfig: {
                  ttsEngine: 'parler',
                  voice: 'Laura',
                  seed: 42
              },
              onProgress: (p) => {
                  const mb = (n) => (n / 1e6).toFixed(1);
                  const line = `▸ Downloading ${p.percentage.toFixed(0)}% (${mb(p.downloaded)}/${mb(p.total)} MB)`;
                  process.stderr.write(process.stderr.isTTY ? `\r${line}` : `${line}\n`);
                  if (p.percentage >= 100)
                      process.stderr.write('\n');
              }
          });
          console.log(`▸ Model loaded: ${modelId}`);
          console.log('▸ Testing Parler Text-to-Speech...');
          const result = textToSpeech({
              modelId,
              text: 'Hey, how are you doing today?',
              inputType: 'text',
              stream: false,
              emotion: 'happy'
          });
          const audioBuffer = await result.buffer;
          const sampleRate = (await result.sampleRate) ?? PARLER_SAMPLE_RATE;
          console.log(`▸ TTS complete. Total samples: ${audioBuffer.length}`);
          console.log('▸ Saving audio to file...');
          createWav(audioBuffer, sampleRate, 'parler-output.wav');
          console.log('▸ Audio saved to parler-output.wav');
          console.log('▸ Playing audio...');
          const audioData = int16ArrayToBuffer(audioBuffer);
          const wavBuffer = Buffer.concat([createWavHeader(audioData.length, sampleRate), audioData]);
          playAudio(wavBuffer);
          console.log('▸ Audio playback complete');
          await unloadModel({ modelId });
          console.log('▸ Model unloaded');
          process.exit(0);
      }
      catch (error) {
          console.error('✖', error);
          process.exit(1);
      }
      ```
    </WrapCode>
  </Tab>

  <Tab value="ts" label="TypeScript">
    <WrapCode>
      ```ts file=<rootDir>/packages/sdk/examples/tts/parler.ts title="tts-parler.ts" lineNumbers
      import {
        loadModel,
        textToSpeech,
        unloadModel,
        type ModelProgressUpdate,
        TTS_MINI_V1_EN_PARLER_TTS_Q8_0
      } from '@qvac/sdk'
      import { createWav, playAudio, int16ArrayToBuffer, createWavHeader } from './utils'

      // Parler-TTS (GGML): description-conditioned speech with per-call voice controls.
      // Uses the registry-hosted Mini v1 Q8_0 model and its native 44.1 kHz output.
      // Only a fallback: the engine reports the rate it actually produced, and
      // `outputSampleRate` (plus the LavaSR enhancer) can move it off this default.
      const PARLER_SAMPLE_RATE = 44100

      try {
        const modelId = await loadModel({
          modelSrc: TTS_MINI_V1_EN_PARLER_TTS_Q8_0,
          modelConfig: {
            ttsEngine: 'parler',
            voice: 'Laura',
            seed: 42
          },
          onProgress: (p: ModelProgressUpdate) => {
            const mb = (n: number) => (n / 1e6).toFixed(1)
            const line = `▸ Downloading ${p.percentage.toFixed(0)}% (${mb(p.downloaded)}/${mb(p.total)} MB)`
            process.stderr.write(process.stderr.isTTY ? `\r${line}` : `${line}\n`)
            if (p.percentage >= 100) process.stderr.write('\n')
          }
        })

        console.log(`▸ Model loaded: ${modelId}`)

        console.log('▸ Testing Parler Text-to-Speech...')
        const result = textToSpeech({
          modelId,
          text: 'Hey, how are you doing today?',
          inputType: 'text',
          stream: false,
          emotion: 'happy'
        })

        const audioBuffer = await result.buffer
        const sampleRate = (await result.sampleRate) ?? PARLER_SAMPLE_RATE
        console.log(`▸ TTS complete. Total samples: ${audioBuffer.length}`)

        console.log('▸ Saving audio to file...')
        createWav(audioBuffer, sampleRate, 'parler-output.wav')
        console.log('▸ Audio saved to parler-output.wav')

        console.log('▸ Playing audio...')
        const audioData = int16ArrayToBuffer(audioBuffer)
        const wavBuffer = Buffer.concat([createWavHeader(audioData.length, sampleRate), audioData])
        playAudio(wavBuffer)
        console.log('▸ Audio playback complete')

        await unloadModel({ modelId })
        console.log('▸ Model unloaded')
        process.exit(0)
      } catch (error) {
        console.error('✖', error)
        process.exit(1)
      }
      ```
    </WrapCode>
  </Tab>
</Tabs>

### CosyVoice3

The following script loads the registry-hosted CosyVoice3 model (the LLM GGUF
downloads its flow/HiFT, voice, and tokenizer companions automatically),
applies per-request emotion conditioning, saves the PCM output, and plays it:

<Tabs>
  <Tab value="js" label="JavaScript" default>
    <WrapCode>
      ```js file=<rootDir>/packages/sdk/dist/examples/tts/cosyvoice3.js title="tts-cosyvoice3.js" lineNumbers
      import { loadModel, textToSpeech, unloadModel, TTS_COSYVOICE3_LLM_COSYVOICE_Q8_0 } from '@qvac/sdk';
      import { createWav, playAudio, int16ArrayToBuffer, createWavHeader } from './utils';
      // CosyVoice3 (GGML): zero-shot speech with one conditioning control per
      // synthesis (emotion, pace, or instruct). The LLM GGUF's registry companion
      // set downloads the flow/HiFT models, baked voice and tokenizer alongside it.
      // Native output is 24 kHz.
      // Only a fallback: the engine reports the rate it actually produced, and
      // `outputSampleRate` (plus the LavaSR enhancer) can move it off this default.
      const COSYVOICE3_SAMPLE_RATE = 24000;
      try {
          const modelId = await loadModel({
              modelSrc: TTS_COSYVOICE3_LLM_COSYVOICE_Q8_0,
              modelConfig: {
                  ttsEngine: 'cosyvoice3',
                  seed: 42
              },
              onProgress: (p) => {
                  const mb = (n) => (n / 1e6).toFixed(1);
                  const line = `▸ Downloading ${p.percentage.toFixed(0)}% (${mb(p.downloaded)}/${mb(p.total)} MB)`;
                  process.stderr.write(process.stderr.isTTY ? `\r${line}` : `${line}\n`);
                  if (p.percentage >= 100)
                      process.stderr.write('\n');
              }
          });
          console.log(`▸ Model loaded: ${modelId}`);
          console.log('▸ Testing CosyVoice3 Text-to-Speech...');
          const result = textToSpeech({
              modelId,
              text: 'Hey, how are you doing today?',
              inputType: 'text',
              stream: false,
              emotion: 'happy'
          });
          const audioBuffer = await result.buffer;
          const sampleRate = (await result.sampleRate) ?? COSYVOICE3_SAMPLE_RATE;
          console.log(`▸ TTS complete. Total samples: ${audioBuffer.length}`);
          console.log('▸ Saving audio to file...');
          createWav(audioBuffer, sampleRate, 'cosyvoice3-output.wav');
          console.log('▸ Audio saved to cosyvoice3-output.wav');
          console.log('▸ Playing audio...');
          const audioData = int16ArrayToBuffer(audioBuffer);
          const wavBuffer = Buffer.concat([createWavHeader(audioData.length, sampleRate), audioData]);
          playAudio(wavBuffer);
          console.log('▸ Audio playback complete');
          await unloadModel({ modelId });
          console.log('▸ Model unloaded');
          process.exit(0);
      }
      catch (error) {
          console.error('✖', error);
          process.exit(1);
      }
      ```
    </WrapCode>
  </Tab>

  <Tab value="ts" label="TypeScript">
    <WrapCode>
      ```ts file=<rootDir>/packages/sdk/examples/tts/cosyvoice3.ts title="tts-cosyvoice3.ts" lineNumbers
      import {
        loadModel,
        textToSpeech,
        unloadModel,
        type ModelProgressUpdate,
        TTS_COSYVOICE3_LLM_COSYVOICE_Q8_0
      } from '@qvac/sdk'
      import { createWav, playAudio, int16ArrayToBuffer, createWavHeader } from './utils'

      // CosyVoice3 (GGML): zero-shot speech with one conditioning control per
      // synthesis (emotion, pace, or instruct). The LLM GGUF's registry companion
      // set downloads the flow/HiFT models, baked voice and tokenizer alongside it.
      // Native output is 24 kHz.
      // Only a fallback: the engine reports the rate it actually produced, and
      // `outputSampleRate` (plus the LavaSR enhancer) can move it off this default.
      const COSYVOICE3_SAMPLE_RATE = 24000

      try {
        const modelId = await loadModel({
          modelSrc: TTS_COSYVOICE3_LLM_COSYVOICE_Q8_0,
          modelConfig: {
            ttsEngine: 'cosyvoice3',
            seed: 42
          },
          onProgress: (p: ModelProgressUpdate) => {
            const mb = (n: number) => (n / 1e6).toFixed(1)
            const line = `▸ Downloading ${p.percentage.toFixed(0)}% (${mb(p.downloaded)}/${mb(p.total)} MB)`
            process.stderr.write(process.stderr.isTTY ? `\r${line}` : `${line}\n`)
            if (p.percentage >= 100) process.stderr.write('\n')
          }
        })

        console.log(`▸ Model loaded: ${modelId}`)

        console.log('▸ Testing CosyVoice3 Text-to-Speech...')
        const result = textToSpeech({
          modelId,
          text: 'Hey, how are you doing today?',
          inputType: 'text',
          stream: false,
          emotion: 'happy'
        })

        const audioBuffer = await result.buffer
        const sampleRate = (await result.sampleRate) ?? COSYVOICE3_SAMPLE_RATE
        console.log(`▸ TTS complete. Total samples: ${audioBuffer.length}`)

        console.log('▸ Saving audio to file...')
        createWav(audioBuffer, sampleRate, 'cosyvoice3-output.wav')
        console.log('▸ Audio saved to cosyvoice3-output.wav')

        console.log('▸ Playing audio...')
        const audioData = int16ArrayToBuffer(audioBuffer)
        const wavBuffer = Buffer.concat([createWavHeader(audioData.length, sampleRate), audioData])
        playAudio(wavBuffer)
        console.log('▸ Audio playback complete')

        await unloadModel({ modelId })
        console.log('▸ Model unloaded')
        process.exit(0)
      } catch (error) {
        console.error('✖', error)
        process.exit(1)
      }
      ```
    </WrapCode>
  </Tab>
</Tabs>

### Audio8

The following script loads the registry-hosted Audio8 LM together with its
codec decoder, synthesizes 44.1 kHz audio with the default sampling settings,
saves the PCM output, and plays it:

<Tabs>
  <Tab value="js" label="JavaScript" default>
    <WrapCode>
      ```js file=<rootDir>/packages/sdk/dist/examples/tts/audio8.js title="tts-audio8.js" lineNumbers
      import { loadModel, textToSpeech, unloadModel, TTS_LM_MULTILINGUAL_AUDIO8_Q8_0, TTS_CODEC_DECODER_AUDIO8_Q8_0 } from '@qvac/sdk';
      import { createWav, playAudio, int16ArrayToBuffer, createWavHeader } from './utils';
      // Audio8 (GGML): DualAR LM + codec decoder, native 44.1 kHz output. The
      // primary modelSrc is the LM GGUF; the codec decoder loads via modelConfig.
      // For zero-shot voice cloning also pass audio8CodecEncoderModelSrc plus
      // referenceAudioSrc/referenceText (a recording and its exact transcript).
      // On macOS and iOS the decoder downloads with its Core ML bundle, which runs
      // the codec's synthesis stack; `stats.codecOnCoreml` reports whether it did.
      // Only a fallback: the engine reports the rate it actually produced, and
      // `outputSampleRate` (plus the LavaSR enhancer) can move it off this default.
      const AUDIO8_SAMPLE_RATE = 44100;
      try {
          const modelId = await loadModel({
              modelSrc: TTS_LM_MULTILINGUAL_AUDIO8_Q8_0,
              modelConfig: {
                  ttsEngine: 'audio8',
                  audio8CodecDecoderModelSrc: TTS_CODEC_DECODER_AUDIO8_Q8_0,
                  temperature: 0.7,
                  topP: 0.9,
                  seed: 42
              },
              onProgress: (p) => {
                  const mb = (n) => (n / 1e6).toFixed(1);
                  const line = `▸ Downloading ${p.percentage.toFixed(0)}% (${mb(p.downloaded)}/${mb(p.total)} MB)`;
                  process.stderr.write(process.stderr.isTTY ? `\r${line}` : `${line}\n`);
                  if (p.percentage >= 100)
                      process.stderr.write('\n');
              }
          });
          console.log(`▸ Model loaded: ${modelId}`);
          console.log('▸ Testing Audio8 Text-to-Speech...');
          const result = textToSpeech({
              modelId,
              text: 'Hey, how are you doing today?',
              inputType: 'text',
              stream: false
          });
          const audioBuffer = await result.buffer;
          const sampleRate = (await result.sampleRate) ?? AUDIO8_SAMPLE_RATE;
          const stats = await result.stats;
          console.log(`▸ TTS complete. Total samples: ${audioBuffer.length}`);
          console.log(`▸ Codec ran on ${stats?.codecOnCoreml === 1 ? 'Core ML' : 'ggml'}`);
          console.log('▸ Saving audio to file...');
          createWav(audioBuffer, sampleRate, 'audio8-output.wav');
          console.log('▸ Audio saved to audio8-output.wav');
          console.log('▸ Playing audio...');
          const audioData = int16ArrayToBuffer(audioBuffer);
          const wavBuffer = Buffer.concat([createWavHeader(audioData.length, sampleRate), audioData]);
          playAudio(wavBuffer);
          console.log('▸ Audio playback complete');
          await unloadModel({ modelId });
          console.log('▸ Model unloaded');
          process.exit(0);
      }
      catch (error) {
          console.error('✖', error);
          process.exit(1);
      }
      ```
    </WrapCode>
  </Tab>

  <Tab value="ts" label="TypeScript">
    <WrapCode>
      ```ts file=<rootDir>/packages/sdk/examples/tts/audio8.ts title="tts-audio8.ts" lineNumbers
      import {
        loadModel,
        textToSpeech,
        unloadModel,
        type ModelProgressUpdate,
        TTS_LM_MULTILINGUAL_AUDIO8_Q8_0,
        TTS_CODEC_DECODER_AUDIO8_Q8_0
      } from '@qvac/sdk'
      import { createWav, playAudio, int16ArrayToBuffer, createWavHeader } from './utils'

      // Audio8 (GGML): DualAR LM + codec decoder, native 44.1 kHz output. The
      // primary modelSrc is the LM GGUF; the codec decoder loads via modelConfig.
      // For zero-shot voice cloning also pass audio8CodecEncoderModelSrc plus
      // referenceAudioSrc/referenceText (a recording and its exact transcript).
      // On macOS and iOS the decoder downloads with its Core ML bundle, which runs
      // the codec's synthesis stack; `stats.codecOnCoreml` reports whether it did.
      // Only a fallback: the engine reports the rate it actually produced, and
      // `outputSampleRate` (plus the LavaSR enhancer) can move it off this default.
      const AUDIO8_SAMPLE_RATE = 44100

      try {
        const modelId = await loadModel({
          modelSrc: TTS_LM_MULTILINGUAL_AUDIO8_Q8_0,
          modelConfig: {
            ttsEngine: 'audio8',
            audio8CodecDecoderModelSrc: TTS_CODEC_DECODER_AUDIO8_Q8_0,
            temperature: 0.7,
            topP: 0.9,
            seed: 42
          },
          onProgress: (p: ModelProgressUpdate) => {
            const mb = (n: number) => (n / 1e6).toFixed(1)
            const line = `▸ Downloading ${p.percentage.toFixed(0)}% (${mb(p.downloaded)}/${mb(p.total)} MB)`
            process.stderr.write(process.stderr.isTTY ? `\r${line}` : `${line}\n`)
            if (p.percentage >= 100) process.stderr.write('\n')
          }
        })

        console.log(`▸ Model loaded: ${modelId}`)

        console.log('▸ Testing Audio8 Text-to-Speech...')
        const result = textToSpeech({
          modelId,
          text: 'Hey, how are you doing today?',
          inputType: 'text',
          stream: false
        })

        const audioBuffer = await result.buffer
        const sampleRate = (await result.sampleRate) ?? AUDIO8_SAMPLE_RATE
        const stats = await result.stats
        console.log(`▸ TTS complete. Total samples: ${audioBuffer.length}`)
        console.log(`▸ Codec ran on ${stats?.codecOnCoreml === 1 ? 'Core ML' : 'ggml'}`)

        console.log('▸ Saving audio to file...')
        createWav(audioBuffer, sampleRate, 'audio8-output.wav')
        console.log('▸ Audio saved to audio8-output.wav')

        console.log('▸ Playing audio...')
        const audioData = int16ArrayToBuffer(audioBuffer)
        const wavBuffer = Buffer.concat([createWavHeader(audioData.length, sampleRate), audioData])
        playAudio(wavBuffer)
        console.log('▸ Audio playback complete')

        await unloadModel({ modelId })
        console.log('▸ Model unloaded')
        process.exit(0)
      } catch (error) {
        console.error('✖', error)
        process.exit(1)
      }
      ```
    </WrapCode>
  </Tab>
</Tabs>

### MOSS

The first script loads the registry-hosted MOSS-TTS backbone and codec,
streams single-speaker speech — optionally in a voice cloned from a 24 kHz
recording — and reports the time to first audio. The second synthesizes a
two-speaker MOSS-TTSD dialogue from a local backbone and the registry codec.
Both save the PCM output with
[`utils.js` / `utils.ts`](#utils):

<Tabs>
  <Tab value="js" label="JavaScript" default>
    <WrapCode>
      ```js file=<rootDir>/packages/sdk/dist/examples/tts/moss.js title="tts-moss.js" lineNumbers
      import { loadModel, textToSpeech, unloadModel, TTS_DELAY_LLM_MOSS_TTS_F16, TTS_CODEC_DECODER_MOSS_TTS_F16, TTS_CODEC_ENCODER_MOSS_TTS_F16 } from '@qvac/sdk';
      import { createWav, playAudio, int16ArrayToBuffer, createWavHeader } from './utils';
      // MOSS (GGML): OpenMOSS MOSS-TTS v1.5 Delay, an 8B backbone (about 17 GB)
      // plus a codec, so desktop only. The primary modelSrc is the backbone; the
      // codec decoder loads via modelConfig, and the codec encoder is only needed to
      // clone a voice from a 24 kHz reference recording (no transcript needed).
      //
      // MOSS streams natively: with `streamChunkTokens` set, `stream: true` emits a
      // chunk every that-many codec frames (12.5 per second) while the backbone is
      // still generating, and the text is synthesized as one utterance, so the
      // `[pause 1.0s]` marker below and `durationTokens` apply to all of it.
      //
      // Uses registry model constants — downloads automatically from QVAC Registry.
      // Usage: node moss.ts [referenceAudio.wav]
      const [referenceAudioSrc] = process.argv.slice(2);
      // Only a fallback: the engine reports the rate it actually produced.
      const MOSS_SAMPLE_RATE = 24000;
      try {
          const modelId = await loadModel({
              modelSrc: TTS_DELAY_LLM_MOSS_TTS_F16,
              modelConfig: {
                  ttsEngine: 'moss',
                  language: 'en',
                  mossCodecDecoderModelSrc: TTS_CODEC_DECODER_MOSS_TTS_F16,
                  ...(referenceAudioSrc
                      ? { mossCodecEncoderModelSrc: TTS_CODEC_ENCODER_MOSS_TTS_F16, referenceAudioSrc }
                      : {}),
                  streamChunkTokens: 25,
                  useGPU: true,
                  seed: 7
              },
              onProgress: (p) => {
                  const mb = (n) => (n / 1e6).toFixed(1);
                  const line = `▸ Downloading ${p.percentage.toFixed(0)}% (${mb(p.downloaded)}/${mb(p.total)} MB)`;
                  process.stderr.write(process.stderr.isTTY ? `\r${line}` : `${line}\n`);
                  if (p.percentage >= 100)
                      process.stderr.write('\n');
              }
          });
          console.log(`▸ Model loaded: ${modelId}`);
          console.log('▸ Streaming MOSS Text-to-Speech...');
          const started = Date.now();
          const result = textToSpeech({
              modelId,
              text: 'Hold on [pause 1.0s] here it comes: speech that starts playing while the rest is still being generated.',
              stream: true
          });
          const samples = [];
          for await (const sample of result.bufferStream) {
              if (samples.length === 0)
                  console.log(`▸ First audio after ${Date.now() - started} ms`);
              samples.push(sample);
          }
          const sampleRate = (await result.sampleRate) ?? MOSS_SAMPLE_RATE;
          const stats = await result.stats;
          console.log(`▸ TTS complete: ${samples.length} samples, ${stats?.generatedFrames ?? '?'} codec frames, RTF ${stats?.realTimeFactor?.toFixed(2) ?? '?'}`);
          createWav(samples, sampleRate, 'moss-output.wav');
          console.log('▸ Audio saved to moss-output.wav');
          console.log('▸ Playing audio...');
          const audioData = int16ArrayToBuffer(samples);
          const wavBuffer = Buffer.concat([createWavHeader(audioData.length, sampleRate), audioData]);
          playAudio(wavBuffer);
          console.log('▸ Audio playback complete');
          await unloadModel({ modelId });
          console.log('▸ Model unloaded');
          process.exit(0);
      }
      catch (error) {
          console.error('✖', error);
          process.exit(1);
      }
      ```
    </WrapCode>
  </Tab>

  <Tab value="ts" label="TypeScript">
    <WrapCode>
      ```ts file=<rootDir>/packages/sdk/examples/tts/moss.ts title="tts-moss.ts" lineNumbers
      import {
        loadModel,
        textToSpeech,
        unloadModel,
        type ModelProgressUpdate,
        TTS_DELAY_LLM_MOSS_TTS_F16,
        TTS_CODEC_DECODER_MOSS_TTS_F16,
        TTS_CODEC_ENCODER_MOSS_TTS_F16
      } from '@qvac/sdk'
      import { createWav, playAudio, int16ArrayToBuffer, createWavHeader } from './utils'

      // MOSS (GGML): OpenMOSS MOSS-TTS v1.5 Delay, an 8B backbone (about 17 GB)
      // plus a codec, so desktop only. The primary modelSrc is the backbone; the
      // codec decoder loads via modelConfig, and the codec encoder is only needed to
      // clone a voice from a 24 kHz reference recording (no transcript needed).
      //
      // MOSS streams natively: with `streamChunkTokens` set, `stream: true` emits a
      // chunk every that-many codec frames (12.5 per second) while the backbone is
      // still generating, and the text is synthesized as one utterance, so the
      // `[pause 1.0s]` marker below and `durationTokens` apply to all of it.
      //
      // Uses registry model constants — downloads automatically from QVAC Registry.
      // Usage: node moss.ts [referenceAudio.wav]
      const [referenceAudioSrc] = process.argv.slice(2)

      // Only a fallback: the engine reports the rate it actually produced.
      const MOSS_SAMPLE_RATE = 24000

      try {
        const modelId = await loadModel({
          modelSrc: TTS_DELAY_LLM_MOSS_TTS_F16,
          modelConfig: {
            ttsEngine: 'moss',
            language: 'en',
            mossCodecDecoderModelSrc: TTS_CODEC_DECODER_MOSS_TTS_F16,
            ...(referenceAudioSrc
              ? { mossCodecEncoderModelSrc: TTS_CODEC_ENCODER_MOSS_TTS_F16, referenceAudioSrc }
              : {}),
            streamChunkTokens: 25,
            useGPU: true,
            seed: 7
          },
          onProgress: (p: ModelProgressUpdate) => {
            const mb = (n: number) => (n / 1e6).toFixed(1)
            const line = `▸ Downloading ${p.percentage.toFixed(0)}% (${mb(p.downloaded)}/${mb(p.total)} MB)`
            process.stderr.write(process.stderr.isTTY ? `\r${line}` : `${line}\n`)
            if (p.percentage >= 100) process.stderr.write('\n')
          }
        })

        console.log(`▸ Model loaded: ${modelId}`)

        console.log('▸ Streaming MOSS Text-to-Speech...')
        const started = Date.now()
        const result = textToSpeech({
          modelId,
          text: 'Hold on [pause 1.0s] here it comes: speech that starts playing while the rest is still being generated.',
          stream: true
        })

        const samples: number[] = []
        for await (const sample of result.bufferStream) {
          if (samples.length === 0) console.log(`▸ First audio after ${Date.now() - started} ms`)
          samples.push(sample)
        }

        const sampleRate = (await result.sampleRate) ?? MOSS_SAMPLE_RATE
        const stats = await result.stats
        console.log(
          `▸ TTS complete: ${samples.length} samples, ${stats?.generatedFrames ?? '?'} codec frames, RTF ${stats?.realTimeFactor?.toFixed(2) ?? '?'}`
        )

        createWav(samples, sampleRate, 'moss-output.wav')
        console.log('▸ Audio saved to moss-output.wav')

        console.log('▸ Playing audio...')
        const audioData = int16ArrayToBuffer(samples)
        const wavBuffer = Buffer.concat([createWavHeader(audioData.length, sampleRate), audioData])
        playAudio(wavBuffer)
        console.log('▸ Audio playback complete')

        await unloadModel({ modelId })
        console.log('▸ Model unloaded')
        process.exit(0)
      } catch (error) {
        console.error('✖', error)
        process.exit(1)
      }
      ```
    </WrapCode>
  </Tab>
</Tabs>

<Tabs>
  <Tab value="js" label="JavaScript" default>
    <WrapCode>
      ```js file=<rootDir>/packages/sdk/dist/examples/tts/moss-dialogue.js title="tts-moss-dialogue.js" lineNumbers
      import { loadModel, textToSpeech, unloadModel, TTS_CODEC_DECODER_MOSS_TTS_F16, TTS_CODEC_ENCODER_MOSS_TTS_F16 } from '@qvac/sdk';
      import { createWav, playAudio, int16ArrayToBuffer, createWavHeader } from './utils';
      // MOSS-TTSD (GGML): multi-speaker dialogue in one pass, each voice cloned from
      // a 24 kHz reference recording. The model continues the references, so the
      // text opens with what each recording says under its speaker tag, followed by
      // the lines to generate; only the new lines come out as audio. A dialogue is
      // never split into sentences, so `stream: true` delivers the native chunks of
      // the whole conversation (set `streamChunkTokens` to get them while it is
      // still being generated).
      //
      // The MOSS-TTSD backbone is not in the QVAC Registry yet, so it is a local
      // file; the codec halves it shares with MOSS-TTS download from the registry.
      // Usage: node moss-dialogue.ts <moss-ttsd.gguf> <s1.wav> "<s1 transcript>" <s2.wav> "<s2 transcript>"
      const [backbone, s1Audio, s1Text, s2Audio, s2Text] = process.argv.slice(2);
      if (!backbone || !s1Audio || !s1Text || !s2Audio || !s2Text) {
          console.error('Usage: node moss-dialogue.ts <moss-ttsd.gguf> <s1.wav> "<s1 transcript>" <s2.wav> "<s2 transcript>"');
          process.exit(1);
      }
      // Only a fallback: the engine reports the rate it actually produced.
      const MOSS_SAMPLE_RATE = 24000;
      try {
          const modelId = await loadModel({
              modelSrc: backbone,
              modelType: 'tts-ggml',
              modelConfig: {
                  ttsEngine: 'moss',
                  language: 'en',
                  mossCodecDecoderModelSrc: TTS_CODEC_DECODER_MOSS_TTS_F16,
                  mossCodecEncoderModelSrc: TTS_CODEC_ENCODER_MOSS_TTS_F16,
                  // One recording per speaker, in the order the text tags them.
                  dialogueReferenceSrcs: [s1Audio, s2Audio],
                  streamChunkTokens: 25,
                  useGPU: true
              }
          });
          console.log(`▸ Model loaded: ${modelId}`);
          console.log('▸ Synthesizing the dialogue...');
          const result = textToSpeech({
              modelId,
              text: `[S1] ${s1Text} [S2] ${s2Text} ` +
                  '[S1] Did the build finish? [S2] Yes, every test passed. [S1] Great, ship it.',
              stream: true
          });
          const samples = [];
          for await (const sample of result.bufferStream)
              samples.push(sample);
          const sampleRate = (await result.sampleRate) ?? MOSS_SAMPLE_RATE;
          console.log(`▸ Dialogue complete: ${samples.length} samples`);
          createWav(samples, sampleRate, 'moss-dialogue-output.wav');
          console.log('▸ Audio saved to moss-dialogue-output.wav');
          console.log('▸ Playing audio...');
          const audioData = int16ArrayToBuffer(samples);
          const wavBuffer = Buffer.concat([createWavHeader(audioData.length, sampleRate), audioData]);
          playAudio(wavBuffer);
          console.log('▸ Audio playback complete');
          await unloadModel({ modelId });
          console.log('▸ Model unloaded');
          process.exit(0);
      }
      catch (error) {
          console.error('✖', error);
          process.exit(1);
      }
      ```
    </WrapCode>
  </Tab>

  <Tab value="ts" label="TypeScript">
    <WrapCode>
      ```ts file=<rootDir>/packages/sdk/examples/tts/moss-dialogue.ts title="tts-moss-dialogue.ts" lineNumbers
      import {
        loadModel,
        textToSpeech,
        unloadModel,
        TTS_CODEC_DECODER_MOSS_TTS_F16,
        TTS_CODEC_ENCODER_MOSS_TTS_F16
      } from '@qvac/sdk'
      import { createWav, playAudio, int16ArrayToBuffer, createWavHeader } from './utils'

      // MOSS-TTSD (GGML): multi-speaker dialogue in one pass, each voice cloned from
      // a 24 kHz reference recording. The model continues the references, so the
      // text opens with what each recording says under its speaker tag, followed by
      // the lines to generate; only the new lines come out as audio. A dialogue is
      // never split into sentences, so `stream: true` delivers the native chunks of
      // the whole conversation (set `streamChunkTokens` to get them while it is
      // still being generated).
      //
      // The MOSS-TTSD backbone is not in the QVAC Registry yet, so it is a local
      // file; the codec halves it shares with MOSS-TTS download from the registry.
      // Usage: node moss-dialogue.ts <moss-ttsd.gguf> <s1.wav> "<s1 transcript>" <s2.wav> "<s2 transcript>"
      const [backbone, s1Audio, s1Text, s2Audio, s2Text] = process.argv.slice(2)
      if (!backbone || !s1Audio || !s1Text || !s2Audio || !s2Text) {
        console.error(
          'Usage: node moss-dialogue.ts <moss-ttsd.gguf> <s1.wav> "<s1 transcript>" <s2.wav> "<s2 transcript>"'
        )
        process.exit(1)
      }

      // Only a fallback: the engine reports the rate it actually produced.
      const MOSS_SAMPLE_RATE = 24000

      try {
        const modelId = await loadModel({
          modelSrc: backbone,
          modelType: 'tts-ggml',
          modelConfig: {
            ttsEngine: 'moss',
            language: 'en',
            mossCodecDecoderModelSrc: TTS_CODEC_DECODER_MOSS_TTS_F16,
            mossCodecEncoderModelSrc: TTS_CODEC_ENCODER_MOSS_TTS_F16,
            // One recording per speaker, in the order the text tags them.
            dialogueReferenceSrcs: [s1Audio, s2Audio],
            streamChunkTokens: 25,
            useGPU: true
          }
        })

        console.log(`▸ Model loaded: ${modelId}`)

        console.log('▸ Synthesizing the dialogue...')
        const result = textToSpeech({
          modelId,
          text:
            `[S1] ${s1Text} [S2] ${s2Text} ` +
            '[S1] Did the build finish? [S2] Yes, every test passed. [S1] Great, ship it.',
          stream: true
        })

        const samples: number[] = []
        for await (const sample of result.bufferStream) samples.push(sample)

        const sampleRate = (await result.sampleRate) ?? MOSS_SAMPLE_RATE
        console.log(`▸ Dialogue complete: ${samples.length} samples`)

        createWav(samples, sampleRate, 'moss-dialogue-output.wav')
        console.log('▸ Audio saved to moss-dialogue-output.wav')

        console.log('▸ Playing audio...')
        const audioData = int16ArrayToBuffer(samples)
        const wavBuffer = Buffer.concat([createWavHeader(audioData.length, sampleRate), audioData])
        playAudio(wavBuffer)
        console.log('▸ Audio playback complete')

        await unloadModel({ modelId })
        console.log('▸ Model unloaded')
        process.exit(0)
      } catch (error) {
        console.error('✖', error)
        process.exit(1)
      }
      ```
    </WrapCode>
  </Tab>
</Tabs>

### Chatterbox with LavaSR enhancer

The following script shows Chatterbox TTS with the LavaSR enhancer, which neurally bandwidth-extends the output to 48 kHz. Use it with [`utils.js` / `utils.ts`](#utils):

<Tabs>
  <Tab value="js" label="JavaScript" default>
    <WrapCode>
      ```js file=<rootDir>/packages/sdk/dist/examples/tts/chatterbox-enhanced.js title="tts-chatterbox-enhanced.js" lineNumbers
      import { loadModel, textToSpeech, unloadModel, TTS_T3_TURBO_EN_CHATTERBOX_Q8_0, TTS_S3GEN_EN_CHATTERBOX, TTS_ENHANCER_LAVASR_FP16 } from '@qvac/sdk';
      import { createWav, playAudio, int16ArrayToBuffer, createWavHeader } from './utils';
      // Chatterbox TTS (GGML) with the LavaSR enhancer: synthesized audio is neurally
      // bandwidth-extended to 48 kHz. Supplying the enhancer GGUF is what enables
      // enhancement — there is no on/off flag — and it forces the output to 48 kHz.
      // Usage: node chatterbox-enhanced.ts [referenceAudioSrc]
      const [referenceAudioSrc] = process.argv.slice(2);
      // Only a fallback; the engine reports the rate it actually produced.
      const ENHANCED_SAMPLE_RATE = 48000;
      try {
          const modelId = await loadModel({
              modelSrc: TTS_T3_TURBO_EN_CHATTERBOX_Q8_0,
              modelConfig: {
                  ttsEngine: 'chatterbox',
                  language: 'en',
                  s3genModelSrc: TTS_S3GEN_EN_CHATTERBOX.src,
                  lavasrEnhancerModelSrc: TTS_ENHANCER_LAVASR_FP16.src,
                  ...(referenceAudioSrc ? { referenceAudioSrc } : {})
              },
              onProgress: (p) => {
                  const mb = (n) => (n / 1e6).toFixed(1);
                  const line = `▸ Downloading ${p.percentage.toFixed(0)}% (${mb(p.downloaded)}/${mb(p.total)} MB)`;
                  process.stderr.write(process.stderr.isTTY ? `\r${line}` : `${line}\n`);
                  if (p.percentage >= 100)
                      process.stderr.write('\n');
              }
          });
          console.log(`▸ Model loaded: ${modelId}`);
          console.log('▸ Testing Text-to-Speech (LavaSR enhancer)...');
          const result = textToSpeech({
              modelId,
              text: `QVAC SDK is the canonical entry point to QVAC. Written in TypeScript, it provides all QVAC capabilities through a unified interface while also abstracting away the complexity of running your application in a JS environment other than Bare. Supported JS environments include Bare, Node.js, Expo and Bun.`,
              inputType: 'text',
              stream: false
          });
          const audioBuffer = await result.buffer;
          const sampleRate = (await result.sampleRate) ?? ENHANCED_SAMPLE_RATE;
          console.log(`▸ TTS complete. Total samples: ${audioBuffer.length}`);
          console.log('▸ Saving audio to file...');
          createWav(audioBuffer, sampleRate, 'chatterbox-enhanced-output.wav');
          console.log('▸ Audio saved to chatterbox-enhanced-output.wav');
          console.log('▸ Playing audio...');
          const audioData = int16ArrayToBuffer(audioBuffer);
          const wavBuffer = Buffer.concat([createWavHeader(audioData.length, sampleRate), audioData]);
          playAudio(wavBuffer);
          console.log('▸ Audio playback complete');
          await unloadModel({ modelId });
          console.log('▸ Model unloaded');
          process.exit(0);
      }
      catch (error) {
          console.error('✖', error);
          process.exit(1);
      }
      ```
    </WrapCode>
  </Tab>

  <Tab value="ts" label="TypeScript">
    <WrapCode>
      ```ts file=<rootDir>/packages/sdk/examples/tts/chatterbox-enhanced.ts title="tts-chatterbox-enhanced.ts" lineNumbers
      import {
        loadModel,
        textToSpeech,
        unloadModel,
        type ModelProgressUpdate,
        TTS_T3_TURBO_EN_CHATTERBOX_Q8_0,
        TTS_S3GEN_EN_CHATTERBOX,
        TTS_ENHANCER_LAVASR_FP16
      } from '@qvac/sdk'
      import { createWav, playAudio, int16ArrayToBuffer, createWavHeader } from './utils'

      // Chatterbox TTS (GGML) with the LavaSR enhancer: synthesized audio is neurally
      // bandwidth-extended to 48 kHz. Supplying the enhancer GGUF is what enables
      // enhancement — there is no on/off flag — and it forces the output to 48 kHz.
      // Usage: node chatterbox-enhanced.ts [referenceAudioSrc]
      const [referenceAudioSrc] = process.argv.slice(2)

      // Only a fallback; the engine reports the rate it actually produced.
      const ENHANCED_SAMPLE_RATE = 48000

      try {
        const modelId = await loadModel({
          modelSrc: TTS_T3_TURBO_EN_CHATTERBOX_Q8_0,
          modelConfig: {
            ttsEngine: 'chatterbox',
            language: 'en',
            s3genModelSrc: TTS_S3GEN_EN_CHATTERBOX.src,
            lavasrEnhancerModelSrc: TTS_ENHANCER_LAVASR_FP16.src,
            ...(referenceAudioSrc ? { referenceAudioSrc } : {})
          },
          onProgress: (p: ModelProgressUpdate) => {
            const mb = (n: number) => (n / 1e6).toFixed(1)
            const line = `▸ Downloading ${p.percentage.toFixed(0)}% (${mb(p.downloaded)}/${mb(p.total)} MB)`
            process.stderr.write(process.stderr.isTTY ? `\r${line}` : `${line}\n`)
            if (p.percentage >= 100) process.stderr.write('\n')
          }
        })

        console.log(`▸ Model loaded: ${modelId}`)

        console.log('▸ Testing Text-to-Speech (LavaSR enhancer)...')
        const result = textToSpeech({
          modelId,
          text: `QVAC SDK is the canonical entry point to QVAC. Written in TypeScript, it provides all QVAC capabilities through a unified interface while also abstracting away the complexity of running your application in a JS environment other than Bare. Supported JS environments include Bare, Node.js, Expo and Bun.`,
          inputType: 'text',
          stream: false
        })

        const audioBuffer = await result.buffer
        const sampleRate = (await result.sampleRate) ?? ENHANCED_SAMPLE_RATE
        console.log(`▸ TTS complete. Total samples: ${audioBuffer.length}`)

        console.log('▸ Saving audio to file...')
        createWav(audioBuffer, sampleRate, 'chatterbox-enhanced-output.wav')
        console.log('▸ Audio saved to chatterbox-enhanced-output.wav')

        console.log('▸ Playing audio...')
        const audioData = int16ArrayToBuffer(audioBuffer)
        const wavBuffer = Buffer.concat([createWavHeader(audioData.length, sampleRate), audioData])
        playAudio(wavBuffer)
        console.log('▸ Audio playback complete')

        await unloadModel({ modelId })
        console.log('▸ Model unloaded')
        process.exit(0)
      } catch (error) {
        console.error('✖', error)
        process.exit(1)
      }
      ```
    </WrapCode>
  </Tab>
</Tabs>

### Supertonic with LavaSR denoiser + enhancer

The following script shows Supertonic TTS with the full LavaSR pipeline: the denoiser cleans the signal first, then the enhancer bandwidth-extends it to 48 kHz. Use it with [`utils.js` / `utils.ts`](#utils):

<Tabs>
  <Tab value="js" label="JavaScript" default>
    <WrapCode>
      ```js file=<rootDir>/packages/sdk/dist/examples/tts/supertonic-enhanced.js title="tts-supertonic-enhanced.js" lineNumbers
      import { loadModel, textToSpeech, unloadModel, TTS_MULTILINGUAL_SUPERTONIC3_Q8_0, TTS_DENOISER_LAVASR_FP16, TTS_ENHANCER_LAVASR_FP16 } from '@qvac/sdk';
      import { createWav, playAudio, int16ArrayToBuffer, createWavHeader } from './utils';
      // Supertonic 3 TTS (GGML) with LavaSR post-processing: the denoiser cleans the
      // synthesized signal first, then the enhancer bandwidth-extends it to 48 kHz.
      // Supplying the enhancer GGUF is what enables enhancement — there is no on/off
      // flag — and it forces the output to 48 kHz regardless of the engine's native
      // rate.
      // Only a fallback; the engine reports the rate it actually produced.
      const ENHANCED_SAMPLE_RATE = 48000;
      try {
          const modelId = await loadModel({
              modelSrc: TTS_MULTILINGUAL_SUPERTONIC3_Q8_0,
              modelConfig: {
                  ttsEngine: 'supertonic',
                  language: 'en',
                  voice: 'F1',
                  // Denoiser runs first (rate-preserving)…
                  lavasrDenoiserModelSrc: TTS_DENOISER_LAVASR_FP16.src,
                  // …then the enhancer bandwidth-extends to 48 kHz.
                  lavasrEnhancerModelSrc: TTS_ENHANCER_LAVASR_FP16.src
              },
              onProgress: (p) => {
                  const mb = (n) => (n / 1e6).toFixed(1);
                  const line = `▸ Downloading ${p.percentage.toFixed(0)}% (${mb(p.downloaded)}/${mb(p.total)} MB)`;
                  process.stderr.write(process.stderr.isTTY ? `\r${line}` : `${line}\n`);
                  if (p.percentage >= 100)
                      process.stderr.write('\n');
              }
          });
          console.log(`▸ Model loaded: ${modelId}`);
          console.log('▸ Testing Text-to-Speech (LavaSR denoiser + enhancer)...');
          const result = textToSpeech({
              modelId,
              text: `QVAC SDK is the canonical entry point to QVAC. Written in TypeScript, it provides all QVAC capabilities through a unified interface while also abstracting away the complexity of running your application in a JS environment other than Bare. Supported JS environments include Bare, Node.js, Expo and Bun.`,
              inputType: 'text',
              stream: false
          });
          const audioBuffer = await result.buffer;
          const sampleRate = (await result.sampleRate) ?? ENHANCED_SAMPLE_RATE;
          console.log(`▸ TTS complete. Total samples: ${audioBuffer.length}`);
          console.log('▸ Saving audio to file...');
          createWav(audioBuffer, sampleRate, 'supertonic-enhanced-output.wav');
          console.log('▸ Audio saved to supertonic-enhanced-output.wav');
          console.log('▸ Playing audio...');
          const audioData = int16ArrayToBuffer(audioBuffer);
          const wavBuffer = Buffer.concat([createWavHeader(audioData.length, sampleRate), audioData]);
          playAudio(wavBuffer);
          console.log('▸ Audio playback complete');
          await unloadModel({ modelId });
          console.log('▸ Model unloaded');
          process.exit(0);
      }
      catch (error) {
          console.error('✖', error);
          process.exit(1);
      }
      ```
    </WrapCode>
  </Tab>

  <Tab value="ts" label="TypeScript">
    <WrapCode>
      ```ts file=<rootDir>/packages/sdk/examples/tts/supertonic-enhanced.ts title="tts-supertonic-enhanced.ts" lineNumbers
      import {
        loadModel,
        textToSpeech,
        unloadModel,
        type ModelProgressUpdate,
        TTS_MULTILINGUAL_SUPERTONIC3_Q8_0,
        TTS_DENOISER_LAVASR_FP16,
        TTS_ENHANCER_LAVASR_FP16
      } from '@qvac/sdk'
      import { createWav, playAudio, int16ArrayToBuffer, createWavHeader } from './utils'

      // Supertonic 3 TTS (GGML) with LavaSR post-processing: the denoiser cleans the
      // synthesized signal first, then the enhancer bandwidth-extends it to 48 kHz.
      // Supplying the enhancer GGUF is what enables enhancement — there is no on/off
      // flag — and it forces the output to 48 kHz regardless of the engine's native
      // rate.
      // Only a fallback; the engine reports the rate it actually produced.
      const ENHANCED_SAMPLE_RATE = 48000

      try {
        const modelId = await loadModel({
          modelSrc: TTS_MULTILINGUAL_SUPERTONIC3_Q8_0,
          modelConfig: {
            ttsEngine: 'supertonic',
            language: 'en',
            voice: 'F1',
            // Denoiser runs first (rate-preserving)…
            lavasrDenoiserModelSrc: TTS_DENOISER_LAVASR_FP16.src,
            // …then the enhancer bandwidth-extends to 48 kHz.
            lavasrEnhancerModelSrc: TTS_ENHANCER_LAVASR_FP16.src
          },
          onProgress: (p: ModelProgressUpdate) => {
            const mb = (n: number) => (n / 1e6).toFixed(1)
            const line = `▸ Downloading ${p.percentage.toFixed(0)}% (${mb(p.downloaded)}/${mb(p.total)} MB)`
            process.stderr.write(process.stderr.isTTY ? `\r${line}` : `${line}\n`)
            if (p.percentage >= 100) process.stderr.write('\n')
          }
        })

        console.log(`▸ Model loaded: ${modelId}`)

        console.log('▸ Testing Text-to-Speech (LavaSR denoiser + enhancer)...')
        const result = textToSpeech({
          modelId,
          text: `QVAC SDK is the canonical entry point to QVAC. Written in TypeScript, it provides all QVAC capabilities through a unified interface while also abstracting away the complexity of running your application in a JS environment other than Bare. Supported JS environments include Bare, Node.js, Expo and Bun.`,
          inputType: 'text',
          stream: false
        })

        const audioBuffer = await result.buffer
        const sampleRate = (await result.sampleRate) ?? ENHANCED_SAMPLE_RATE
        console.log(`▸ TTS complete. Total samples: ${audioBuffer.length}`)

        console.log('▸ Saving audio to file...')
        createWav(audioBuffer, sampleRate, 'supertonic-enhanced-output.wav')
        console.log('▸ Audio saved to supertonic-enhanced-output.wav')

        console.log('▸ Playing audio...')
        const audioData = int16ArrayToBuffer(audioBuffer)
        const wavBuffer = Buffer.concat([createWavHeader(audioData.length, sampleRate), audioData])
        playAudio(wavBuffer)
        console.log('▸ Audio playback complete')

        await unloadModel({ modelId })
        console.log('▸ Model unloaded')
        process.exit(0)
      } catch (error) {
        console.error('✖', error)
        process.exit(1)
      }
      ```
    </WrapCode>
  </Tab>
</Tabs>

### Utils

The following helper script is used by the examples above to convert the raw PCM samples returned by `textToSpeech()` into a WAV file and play it back:

<Tabs>
  <Tab value="js" label="JavaScript" default>
    <WrapCode>
      ```js file=<rootDir>/packages/sdk/dist/examples/tts/utils.js title="utils.js" lineNumbers
      import { writeFileSync, unlinkSync } from 'fs';
      import { spawn, spawnSync } from 'child_process';
      import { platform, tmpdir } from 'os';
      import { join } from 'path';
      /**
       * Create WAV header for 16-bit PCM audio
       */
      export function createWavHeader(dataLength, sampleRate) {
          const header = Buffer.alloc(44);
          // RIFF header
          header.write('RIFF', 0);
          header.writeUInt32LE(36 + dataLength, 4);
          header.write('WAVE', 8);
          // fmt chunk
          header.write('fmt ', 12);
          header.writeUInt32LE(16, 16); // fmt chunk size
          header.writeUInt16LE(1, 20); // PCM format
          header.writeUInt16LE(1, 22); // mono
          header.writeUInt32LE(sampleRate, 24);
          header.writeUInt32LE(sampleRate * 2, 28); // byte rate
          header.writeUInt16LE(2, 32); // block align
          header.writeUInt16LE(16, 34); // bits per sample
          // data chunk
          header.write('data', 36);
          header.writeUInt32LE(dataLength, 40);
          return header;
      }
      /**
       * Convert Int16Array to Buffer
       */
      export function int16ArrayToBuffer(samples) {
          const buffer = Buffer.alloc(samples.length * 2);
          for (let i = 0; i < samples.length; i++) {
              const value = Math.max(-32768, Math.min(32767, Math.round(samples[i] ?? 0)));
              buffer.writeInt16LE(value, i * 2);
          }
          return buffer;
      }
      /**
       * Create and save WAV file
       */
      export function createWav(audioBuffer, sampleRate, filename) {
          const audioData = int16ArrayToBuffer(audioBuffer);
          const wavHeader = createWavHeader(audioData.length, sampleRate);
          const wavFile = Buffer.concat([wavHeader, audioData]);
          writeFileSync(filename, wavFile);
          console.log(`▸ WAV file saved as: ${filename}`);
      }
      /**
       * Play a WAV buffer by streaming it into ffplay over stdin.
       *
       * ffplay ships with ffmpeg and is cross-platform (macOS/Linux/Windows), so
       * we avoid the old "write to /tmp then shell out to afplay/aplay/powershell"
       * dance — no temp files, no platform switch, no hardcoded /tmp path (which
       * doesn't exist on Windows). Requires ffplay on PATH.
       */
      /**
       * Play one mono s16le PCM chunk (as a minimal WAV) and wait for the player to finish.
       * Chunks are played sequentially when awaited in order — suitable for streaming TTS output.
       */
      export function playPcmInt16Chunk(samples, sampleRate) {
          if (samples.length === 0) {
              return Promise.resolve();
          }
          const audioData = int16ArrayToBuffer(samples);
          const wavHeader = createWavHeader(audioData.length, sampleRate);
          const wavFile = Buffer.concat([wavHeader, audioData]);
          // `os.tmpdir()` resolves to the OS-specific temp directory (e.g. `%TEMP%`
          // on Windows), so the Windows branch below no longer tries to read a
          // POSIX-only `/tmp/...` path.
          const tempFile = join(tmpdir(), `qvac-tts-chunk-${Date.now()}-${Math.random().toString(16).slice(2)}.wav`);
          writeFileSync(tempFile, wavFile);
          const currentPlatform = platform();
          let audioPlayer;
          let args;
          switch (currentPlatform) {
              case 'darwin':
                  audioPlayer = 'afplay';
                  args = [tempFile];
                  break;
              case 'linux':
                  audioPlayer = 'aplay';
                  args = [tempFile];
                  break;
              case 'win32':
                  audioPlayer = 'powershell';
                  args = [
                      '-Command',
                      `Add-Type -AssemblyName presentationCore; (New-Object Media.SoundPlayer).LoadStream([System.IO.File]::ReadAllBytes('${tempFile}')).PlaySync()`
                  ];
                  break;
              default:
                  audioPlayer = 'aplay';
                  args = [tempFile];
          }
          return new Promise(function (resolve, reject) {
              const proc = spawn(audioPlayer, args, { stdio: 'ignore' });
              proc.on('error', function (err) {
                  try {
                      unlinkSync(tempFile);
                  }
                  catch {
                      // ignore
                  }
                  reject(err);
              });
              proc.on('close', function (code) {
                  try {
                      unlinkSync(tempFile);
                  }
                  catch {
                      // ignore
                  }
                  if (code === 0) {
                      resolve();
                  }
                  else {
                      reject(new Error(`Audio player exited with code ${code}`));
                  }
              });
          });
      }
      export function playAudio(audioBuffer) {
          const result = spawnSync('ffplay', ['-hide_banner', '-loglevel', 'error', '-autoexit', '-nodisp', '-i', 'pipe:0'], {
              input: audioBuffer,
              stdio: ['pipe', 'inherit', 'inherit']
          });
          if (result.error) {
              const code = result.error.code;
              if (code === 'ENOENT') {
                  throw new Error('ffplay not found on PATH. Install ffmpeg (ffplay ships with it) and retry.');
              }
              throw new Error(`ffplay failed: ${result.error.message}`);
          }
          if (result.status !== 0) {
              throw new Error(`ffplay exited with code ${result.status}`);
          }
      }
      ```
    </WrapCode>
  </Tab>

  <Tab value="ts" label="TypeScript">
    <WrapCode>
      ```ts file=<rootDir>/packages/sdk/examples/tts/utils.ts title="utils.ts" lineNumbers
      import { writeFileSync, unlinkSync } from 'fs'
      import { spawn, spawnSync } from 'child_process'
      import { platform, tmpdir } from 'os'
      import { join } from 'path'

      /**
       * Create WAV header for 16-bit PCM audio
       */
      export function createWavHeader(dataLength: number, sampleRate: number): Buffer {
        const header = Buffer.alloc(44)

        // RIFF header
        header.write('RIFF', 0)
        header.writeUInt32LE(36 + dataLength, 4)
        header.write('WAVE', 8)

        // fmt chunk
        header.write('fmt ', 12)
        header.writeUInt32LE(16, 16) // fmt chunk size
        header.writeUInt16LE(1, 20) // PCM format
        header.writeUInt16LE(1, 22) // mono
        header.writeUInt32LE(sampleRate, 24)
        header.writeUInt32LE(sampleRate * 2, 28) // byte rate
        header.writeUInt16LE(2, 32) // block align
        header.writeUInt16LE(16, 34) // bits per sample

        // data chunk
        header.write('data', 36)
        header.writeUInt32LE(dataLength, 40)

        return header
      }

      /**
       * Convert Int16Array to Buffer
       */
      export function int16ArrayToBuffer(samples: number[]): Buffer {
        const buffer = Buffer.alloc(samples.length * 2)
        for (let i = 0; i < samples.length; i++) {
          const value = Math.max(-32768, Math.min(32767, Math.round(samples[i] ?? 0)))
          buffer.writeInt16LE(value, i * 2)
        }
        return buffer
      }

      /**
       * Create and save WAV file
       */
      export function createWav(audioBuffer: number[], sampleRate: number, filename: string): void {
        const audioData = int16ArrayToBuffer(audioBuffer)
        const wavHeader = createWavHeader(audioData.length, sampleRate)
        const wavFile = Buffer.concat([wavHeader, audioData])

        writeFileSync(filename, wavFile)
        console.log(`▸ WAV file saved as: ${filename}`)
      }

      /**
       * Play a WAV buffer by streaming it into ffplay over stdin.
       *
       * ffplay ships with ffmpeg and is cross-platform (macOS/Linux/Windows), so
       * we avoid the old "write to /tmp then shell out to afplay/aplay/powershell"
       * dance — no temp files, no platform switch, no hardcoded /tmp path (which
       * doesn't exist on Windows). Requires ffplay on PATH.
       */
      /**
       * Play one mono s16le PCM chunk (as a minimal WAV) and wait for the player to finish.
       * Chunks are played sequentially when awaited in order — suitable for streaming TTS output.
       */
      export function playPcmInt16Chunk(samples: number[], sampleRate: number): Promise<void> {
        if (samples.length === 0) {
          return Promise.resolve()
        }

        const audioData = int16ArrayToBuffer(samples)
        const wavHeader = createWavHeader(audioData.length, sampleRate)
        const wavFile = Buffer.concat([wavHeader, audioData])
        // `os.tmpdir()` resolves to the OS-specific temp directory (e.g. `%TEMP%`
        // on Windows), so the Windows branch below no longer tries to read a
        // POSIX-only `/tmp/...` path.
        const tempFile = join(
          tmpdir(),
          `qvac-tts-chunk-${Date.now()}-${Math.random().toString(16).slice(2)}.wav`
        )
        writeFileSync(tempFile, wavFile)

        const currentPlatform = platform()
        let audioPlayer: string
        let args: string[]

        switch (currentPlatform) {
          case 'darwin':
            audioPlayer = 'afplay'
            args = [tempFile]
            break
          case 'linux':
            audioPlayer = 'aplay'
            args = [tempFile]
            break
          case 'win32':
            audioPlayer = 'powershell'
            args = [
              '-Command',
              `Add-Type -AssemblyName presentationCore; (New-Object Media.SoundPlayer).LoadStream([System.IO.File]::ReadAllBytes('${tempFile}')).PlaySync()`
            ]
            break
          default:
            audioPlayer = 'aplay'
            args = [tempFile]
        }

        return new Promise(function (resolve, reject) {
          const proc = spawn(audioPlayer, args, { stdio: 'ignore' })
          proc.on('error', function (err) {
            try {
              unlinkSync(tempFile)
            } catch {
              // ignore
            }
            reject(err)
          })
          proc.on('close', function (code) {
            try {
              unlinkSync(tempFile)
            } catch {
              // ignore
            }
            if (code === 0) {
              resolve()
            } else {
              reject(new Error(`Audio player exited with code ${code}`))
            }
          })
        })
      }

      export function playAudio(audioBuffer: Buffer): void {
        const result = spawnSync(
          'ffplay',
          ['-hide_banner', '-loglevel', 'error', '-autoexit', '-nodisp', '-i', 'pipe:0'],
          {
            input: audioBuffer,
            stdio: ['pipe', 'inherit', 'inherit']
          }
        )

        if (result.error) {
          const code = (result.error as NodeJS.ErrnoException).code
          if (code === 'ENOENT') {
            throw new Error('ffplay not found on PATH. Install ffmpeg (ffplay ships with it) and retry.')
          }
          throw new Error(`ffplay failed: ${result.error.message}`)
        }
        if (result.status !== 0) {
          throw new Error(`ffplay exited with code ${result.status}`)
        }
      }
      ```
    </WrapCode>
  </Tab>
</Tabs>

<Callout type="success">
  **Tip:** all examples throughout this documentation are self-contained and
  runnable. For instructions on how to run them, see the [JS/TS
  quickstart](/sdk/v0.20/js-ts-sdk#quickstart) or the [Python
  quickstart](/sdk/v0.20/python-sdk#quickstart).
</Callout>
