New: TranslatePsy-AfriSLM translates directly between 19 African languages, offline.
QVAC Logo
SDKTranscription
v0.20, not the current release

Transcription

Automatic speech recognition (ASR) for speech-to-text — i.e., generate text transcriptions from audio input.

Overview

Transcription uses your choice of either qvac-fabric-speech.cpp or NVIDIA Parakeet (via the GGML-based parakeet-cpp engine) as inference engine. Load a model using modelType: "whisper" for qvac-fabric-speech.cpp, or modelType: "parakeet-transcription" for Parakeet. Parakeet supports multilingual transcription (TDT), English-only transcription (CTC), English transcription with a single model that supports both batch and duplex streaming (Unified), Indic-language transcription (Indic Conformer CTC), speaker diarization (Sortformer), and end-of-utterance detection (EOU) for duplex streaming.

Provide audio input as audioChunk, either as a file path (string) or an in-memory audio buffer.

transcribe() returns the full transcription as a single string. Its returned promise also exposes stats: await it after transcription to read terminal engine statistics. For Parakeet offline ASR, encoderOnCoreml reports whether a Core ML sidecar loaded, while encoderUsedCoreml reports whether this run encoded on Core ML. The stats promise resolves to undefined when the engine does not report statistics. If you need partial results as they become available, use transcribeStream() to receive text chunks in real-time. Both whisper and parakeet expose duplex transcribeStream() sessions; see "Streaming with transcribeStream()" below.

Functions

Use the following sequence of function calls:

  1. loadModel()
  2. transcribe() or transcribeStream()
  3. unloadModel()

For how to use each function, see SDK — API reference.

Models

qvac-fabric-speech.cpp

You should load two models:

  • a whisper.cpp-compatible model for transcription. Model file format: *.bin; and
  • a VAD model (e.g., Silero) converted to GGML. Model file format: *.bin (optional, recommended).

Parakeet

Parakeet uses a GGUF checkpoint per variant, and the addon detects TDT / CTC / Unified / Sortformer / EOU from its GGUF metadata. Supply the GGUF via the top-level modelSrc:

await loadModel({
  modelSrc: PARAKEET_TDT_0_6B_V3_Q8_0,    // multilingual, ~750MB
  modelType: "parakeet-transcription",
});

await loadModel({
  modelSrc: PARAKEET_CTC_0_6B_Q8_0,       // english-only, streaming-capable
  modelType: "parakeet-transcription",
});

await loadModel({
  modelSrc: PARAKEET_UNIFIED_0_6B_Q8_0,   // english; batch + streaming, ~740MB
  modelType: "parakeet-transcription",
});

await loadModel({
  modelSrc: PARAKEET_SORTFORMER_4SPK_V2_1_Q8_0,  // 4-speaker diarization
  modelType: "parakeet-transcription",
});

await loadModel({
  modelSrc: PARAKEET_EOU_120M_V1_Q8_0,    // end-of-utterance detection
  modelType: "parakeet-transcription",
});

await loadModel({
  modelSrc: PARAKEET_INDIC_CONFORMER_600M_Q8_0,  // Indic languages
  modelType: "parakeet-transcription",
  modelConfig: { language: "hi" },              // required; e.g. "hi", "ta"
});

On macOS and iOS, loading a supported Parakeet registry constant also downloads its complete Core ML encoder bundle and stages the .mlmodelc directory beside the GGUF. This applies to TDT 0.6B v3, Unified English 0.6B, EOU 120M v1, and streaming Sortformer v2.1. The sidecar weights increase the first download and disk use. Other platforms download the GGUF alone; if a sidecar is unavailable or cannot load, the GGUF encoder is used.

For model artifacts available as constants, see SDK — Models.

Migrating from pre-0.6 Parakeet (ONNX multi-file): the legacy multi-file ONNX modelConfig shape (parakeetEncoderSrc / parakeetDecoderSrc / parakeetVocabSrc / parakeetPreprocessorSrc, plus parakeetCtcModelSrc / parakeetTokenizerSrc and parakeetSortformerSrc for the CTC/Sortformer variants) is no longer supported. Passing any of those fields raises a structured LegacyParakeetModelDeprecatedError with a migration message. The legacy ONNX constants (e.g. PARAKEET_TDT_ENCODER_INT8, PARAKEET_CTC_FP32, PARAKEET_SORTFORMER_FP32) remain exported for one minor cycle for codemod migrations only and will be removed in a future release.

On VAD: when using qvac-fabric-speech.cpp, you can optionally provide a separate model for voice activity detection (VAD); this is recommended. In turn, Parakeet handles VAD internally, so no additional model or configuration is required.

Streaming with transcribeStream()

transcribeStream() opens a duplex session for both engines — write audio chunks via session.write(...), iterate events with for await (const event of session) { ... }. Events are typed as a discriminated union { type }:

  • { type: "text", text } — incremental transcript text.
  • { type: "segment", segment } — segment metadata when metadata: true, on both engines. Parakeet segments also carry isEndOfTurn and startsWord.
  • { type: "vad", speaking, probability, source } — voice-activity-detection state (whisper-only; source is "silero").
  • { type: "endOfTurn", source: "whisper", silenceDurationMs } — turn boundary detected from a measured silence window (whisper).
  • { type: "endOfTurn", source: "parakeet" } — turn boundary detected from the EOU model's <EOU> token (parakeet; no silence window — the event is token-driven).

The source field on endOfTurn lets consumers narrow the union: whisper events always carry a numeric silenceDurationMs; parakeet events never do.

After the iterator completes, await session.stats to read terminal engine statistics. The promise resolves to undefined when the engine does not report stats.

for await (const event of session) {
  // Handle text, segment, VAD, or end-of-turn events.
}

const stats = await session.stats;
console.log(stats?.audioDuration, stats?.realTimeFactor);

Wire compatibility: post-0.6 servers emit source on every endOfTurn frame. SDK parsers still accept the legacy whisper wire shape { silenceDurationMs } (no source) and normalize it to source: "whisper". Upgrade client and server together when using parakeet source: "parakeet" events — older servers never emit that branch.

Parakeet duplex streaming

Pass parakeetStreamingConfig to transcribeStream() to override per-call streaming knobs (each falls back to its parakeetConfig.streaming* load-time counterpart):

const session = await transcribeStream({
  modelId,
  parakeetStreamingConfig: {
    chunkMs: 1000,            // encoder cadence
    historyMs: 30000,         // sortformer rolling-history window
    leftContextMs: 500,       // ASR encoder left-context window
    rightLookaheadMs: 200,    // ASR encoder right-lookahead window
    emitPartials: true,       // emit partial segments before chunk boundaries
    emitEnergyVad: false,     // CTC/TDT energy-based VAD hint (engine-internal)
  },
});

for await (const event of session) {
  switch (event.type) {
    case "text":
      process.stdout.write(event.text);
      break;
    case "endOfTurn":
      // event.source: "whisper" | "parakeet"
      console.log("\n[endOfTurn] turn boundary detected\n");
      break;
  }
}

The synthetic { type: "endOfTurn", source: "parakeet" } event surfaces whenever the EOU model emits an <EOU> token, and is the parakeet equivalent of whisper's silence-window EOU. Pair it with the PARAKEET_EOU_120M_V1_Q8_0 checkpoint when you need explicit turn boundaries from parakeet.

Examples

qvac-fabric-speech.cpp

The following script shows an example of qvac-fabric-speech.cpp transcription with prompt-guided decoding, VAD, and GPU acceleration:

whispercpp-prompt.js
/**
 * Whisper transcription with prompt example.
 *
 * Usage:
 *   bun examples/asr/whispercpp-prompt.ts
 *
 * This example requires a test audio file (default: examples/audio/sample-16khz.wav).
 * Sample audio files are available in the QVAC source repository, but not included in the published npm package.
 * Set audioChunk to a custom WAV, or download the default audio into examples/audio/:
 *   https://github.com/tetherto/qvac/blob/main/packages/sdk/examples/audio/sample-16khz.wav
 */
import { loadModel, unloadModel, transcribe, WHISPER_TINY } from '@qvac/sdk';
try {
    console.log('▸ Starting Whisper transcription with prompt example...');
    // Load the Whisper model
    console.log('▸ Loading Whisper model...');
    const modelId = await loadModel({
        modelSrc: WHISPER_TINY,
        modelConfig: {
            audio_format: 'f32le',
            // Sampling strategy
            strategy: 'greedy',
            n_threads: 4,
            // Transcription options
            language: 'en',
            translate: false,
            no_timestamps: false,
            single_segment: false,
            print_timestamps: true,
            token_timestamps: true,
            // Quality settings
            temperature: 0.0,
            suppress_blank: true,
            suppress_nst: true,
            // Advanced tuning
            entropy_thold: 2.4,
            logprob_thold: -1.0,
            // VAD configuration
            vad_params: {
                threshold: 0.35,
                min_speech_duration_ms: 200,
                min_silence_duration_ms: 150,
                max_speech_duration_s: 30.0,
                speech_pad_ms: 600,
                samples_overlap: 0.3
            },
            // Context parameters for GPU
            contextParams: {
                use_gpu: true,
                flash_attn: true,
                gpu_device: 0
            }
        },
        onProgress: (p) => {
            const mb = (n) => (n / 1e6).toFixed(1);
            const line = `▸ Downloading ${p.percentage.toFixed(0)}% (${mb(p.downloaded)}/${mb(p.total)} MB)`;
            process.stderr.write(process.stderr.isTTY ? `\r${line}` : `${line}\n`);
            if (p.percentage >= 100)
                process.stderr.write('\n');
        }
    });
    console.log(`▸ Whisper model loaded with ID: ${modelId}`);
    // Perform transcription
    console.log('▸ Transcribing audio...');
    const text = await transcribe({
        modelId,
        audioChunk: 'examples/audio/sample-16khz.wav',
        prompt: 'This is a test recording with clear speech and proper punctuation.'
    });
    console.log('▸ Transcription result:');
    console.log(text);
    // Unload the model when done
    console.log('▸ Unloading Whisper model...');
    await unloadModel({ modelId });
    console.log('▸ Whisper model unloaded successfully');
    process.exit(0);
}
catch (error) {
    console.error('✖', error);
    process.exit(1);
}

Parakeet TDT

The following script shows an example of multilingual transcription using the Parakeet TDT model from a WAV file:

parakeet-tdt-filesystem.js
/**
 * Parakeet TDT transcription from a WAV file.
 *
 * Usage:
 *   bun run examples/asr/parakeet-tdt-filesystem.ts <wav-file> [parakeet-tdt-gguf]
 *
 * Loads a single GGUF checkpoint (`PARAKEET_TDT_0_6B_V3_Q8_0` by default) and
 * transcribes the file with the batch `transcribe` API. Omit the model
 * argument to use the registry constant.
 *
 * Audio should be 16 kHz mono PCM in a WAV container.
 */
import { loadModel, unloadModel, transcribe, PARAKEET_TDT_0_6B_V3_Q8_0 } from '@qvac/sdk';
const args = process.argv.slice(2);
if (!args[0]) {
    console.error('Usage: bun run examples/asr/parakeet-tdt-filesystem.ts <wav-file-path> ' +
        '[parakeet-tdt-gguf]');
    console.error('\nIf the model path is omitted, defaults to the registry model.');
    process.exit(1);
}
const audioFilePath = args[0];
const parakeetModelSrc = args[1] ?? PARAKEET_TDT_0_6B_V3_Q8_0;
try {
    console.log('▸ Starting Parakeet transcription example...');
    console.log('▸ Loading Parakeet model...');
    const modelId = await loadModel({
        modelSrc: parakeetModelSrc,
        modelType: 'parakeet-transcription',
        onProgress: (p) => {
            const mb = (n) => (n / 1e6).toFixed(1);
            const line = `▸ Downloading ${p.percentage.toFixed(0)}% (${mb(p.downloaded)}/${mb(p.total)} MB)`;
            process.stderr.write(process.stderr.isTTY ? `\r${line}` : `${line}\n`);
            if (p.percentage >= 100)
                process.stderr.write('\n');
        }
    });
    console.log(`▸ Parakeet model loaded with ID: ${modelId}`);
    console.log('▸ Transcribing audio...');
    const text = await transcribe({ modelId, audioChunk: audioFilePath });
    console.log(text);
    console.log('▸ Unloading Parakeet model...');
    await unloadModel({ modelId });
    console.log('▸ Parakeet model unloaded successfully');
}
catch (error) {
    console.error('✖', error);
    process.exit(1);
}

Parakeet CTC

The following script shows an example of English-only transcription using the Parakeet CTC model from a WAV file:

parakeet-ctc-filesystem.js
/**
 * Parakeet CTC transcription from a WAV file.
 *
 * Usage:
 *   bun run examples/asr/parakeet-ctc-filesystem.ts <wav-file> [parakeet-ctc-gguf]
 *
 * Loads a single GGUF checkpoint (`PARAKEET_CTC_0_6B_Q8_0` by default) and
 * transcribes the file with the batch `transcribe` API. Omit the model
 * argument to use the registry constant.
 *
 * Audio should be 16 kHz mono PCM in a WAV container.
 */
import { loadModel, unloadModel, transcribe, PARAKEET_CTC_0_6B_Q8_0 } from '@qvac/sdk';
const args = process.argv.slice(2);
if (!args[0]) {
    console.error('Usage: bun run examples/asr/parakeet-ctc-filesystem.ts <wav-file> ' + '[parakeet-ctc-gguf]');
    console.error('\nIf the model path is omitted, defaults to the registry model.');
    process.exit(1);
}
const audioFilePath = args[0];
const parakeetModelSrc = args[1] ?? PARAKEET_CTC_0_6B_Q8_0;
try {
    console.log('▸ Loading Parakeet CTC model...');
    const modelId = await loadModel({
        modelSrc: parakeetModelSrc,
        modelType: 'parakeet-transcription',
        onProgress: (p) => {
            const mb = (n) => (n / 1e6).toFixed(1);
            const line = `▸ Downloading ${p.percentage.toFixed(0)}% (${mb(p.downloaded)}/${mb(p.total)} MB)`;
            process.stderr.write(process.stderr.isTTY ? `\r${line}` : `${line}\n`);
            if (p.percentage >= 100)
                process.stderr.write('\n');
        }
    });
    console.log(`▸ Parakeet CTC model loaded with ID: ${modelId}`);
    console.log('▸ Transcribing audio...');
    const text = await transcribe({ modelId, audioChunk: audioFilePath });
    console.log(text);
    console.log('▸ Unloading model...');
    await unloadModel({ modelId });
    console.log('▸ Done');
}
catch (error) {
    console.error('✖', error);
    process.exit(1);
}

Parakeet Unified

The following script shows an example of English transcription using the Parakeet Unified model from a WAV file. Load it once and use it with either transcribe() for batch or transcribeStream() for low-latency streaming.

parakeet-unified-filesystem.js
/**
 * Parakeet Unified transcription from a WAV file.
 *
 * Usage:
 *   bun run examples/asr/parakeet-unified-filesystem.ts <wav-file> [parakeet-unified-gguf]
 *
 * Loads a single GGUF checkpoint (`PARAKEET_UNIFIED_0_6B_Q8_0` by default) and
 * transcribes the file with the batch `transcribe` API. The Unified RNN-T
 * checkpoint is English-only and serves both batch and low-latency streaming
 * from the same GGUF; the engine auto-detects the model type from the GGUF
 * metadata. Omit the model argument to use the registry constant.
 *
 * Audio should be 16 kHz mono PCM in a WAV container.
 */
import { loadModel, unloadModel, transcribe, PARAKEET_UNIFIED_0_6B_Q8_0 } from '@qvac/sdk';
const args = process.argv.slice(2);
if (!args[0]) {
    console.error('Usage: bun run examples/asr/parakeet-unified-filesystem.ts <wav-file-path> ' +
        '[parakeet-unified-gguf]');
    console.error('\nIf the model path is omitted, defaults to the registry model.');
    process.exit(1);
}
const audioFilePath = args[0];
const parakeetModelSrc = args[1] ?? PARAKEET_UNIFIED_0_6B_Q8_0;
try {
    console.log('▸ Starting Parakeet Unified transcription example...');
    console.log('▸ Loading Parakeet Unified model...');
    const modelId = await loadModel({
        modelSrc: parakeetModelSrc,
        modelType: 'parakeet-transcription',
        onProgress: (p) => {
            const mb = (n) => (n / 1e6).toFixed(1);
            const line = `▸ Downloading ${p.percentage.toFixed(0)}% (${mb(p.downloaded)}/${mb(p.total)} MB)`;
            process.stderr.write(process.stderr.isTTY ? `\r${line}` : `${line}\n`);
            if (p.percentage >= 100)
                process.stderr.write('\n');
        }
    });
    console.log(`▸ Parakeet Unified model loaded with ID: ${modelId}`);
    console.log('▸ Transcribing audio...');
    const text = await transcribe({ modelId, audioChunk: audioFilePath });
    console.log(text);
    console.log('▸ Unloading Parakeet Unified model...');
    await unloadModel({ modelId });
    console.log('▸ Parakeet Unified model unloaded successfully');
}
catch (error) {
    console.error('✖', error);
    process.exit(1);
}

Parakeet Unified with Core ML

The following macOS script loads a regenerated Unified GGUF from disk, downloads its five Core ML encoder files through the registry, and transcribes a WAV file:

parakeet-unified-coreml.js
/**
 * Transcribe a WAV file with Parakeet Unified and its Core ML encoder sidecar.
 *
 * From packages/sdk:
 *   bun run examples/asr/parakeet-unified-coreml.ts \
 *     /path/to/parakeet-unified-en-0.6b.q4_0.gguf [16-kHz-wav-file]
 *
 * The GGUF must be a regenerated Unified artifact. The five compiled Core ML
 * files are downloaded through the SDK registry and staged beside the GGUF.
 * The default audio is examples/audio/sample-16khz.wav.
 */
import { downloadAsset, getModelInfo, loadModel, transcribe, unloadModel } from '@qvac/sdk';
import { link, mkdir, mkdtemp, open, rm, symlink } from 'node:fs/promises';
import { tmpdir } from 'node:os';
import { basename, join, resolve } from 'node:path';
import { fileURLToPath } from 'node:url';
const MODEL_STEM = 'parakeet-unified-en-0.6b';
const COREML_FILES = [
    ['PARAKEET_UNIFIED_COREMLDATA', 'analytics/coremldata.bin'],
    ['PARAKEET_UNIFIED_COREMLDATA_1', 'coremldata.bin'],
    ['PARAKEET_UNIFIED_METADATA', 'metadata.json'],
    ['PARAKEET_UNIFIED_MODEL', 'model.mil'],
    ['PARAKEET_UNIFIED_WEIGHT', 'weights/weight.bin']
];
const defaultAudioPath = fileURLToPath(new URL('../audio/sample-16khz.wav', import.meta.url));
const modelPath = process.argv[2];
const audioPath = process.argv[3] ?? defaultAudioPath;
async function checkModel(path) {
    const name = basename(path);
    if (!/^parakeet-unified-en-0\.6b\.(?:f16|q4_0|q8_0)\.gguf$/.test(name)) {
        throw new Error(`Expected a regenerated Unified GGUF, got: ${name}`);
    }
    const file = await open(path, 'r');
    try {
        const header = Buffer.alloc(1024 * 1024);
        const { bytesRead } = await file.read(header, 0, header.length, 0);
        if (!header.subarray(0, bytesRead).includes('parakeet.unified.left_context_frames')) {
            throw new Error(`Unified GGUF lacks the new streaming metadata: ${path}`);
        }
    }
    finally {
        await file.close();
    }
}
async function stageModel(path) {
    if (process.platform !== 'darwin')
        throw new Error('Core ML requires macOS');
    await checkModel(path);
    const dir = await mkdtemp(join(tmpdir(), 'qvac-unified-coreml-'));
    try {
        const modelSrc = join(dir, basename(path));
        await symlink(resolve(path), modelSrc);
        const bundleDir = join(dir, `${MODEL_STEM}-encoder.mlmodelc`);
        for (const [name, relativePath] of COREML_FILES) {
            const info = await getModelInfo({ name });
            if (!info.isCached) {
                if (!info.registrySource || !info.registryPath) {
                    throw new Error(`Registry location missing for ${name}`);
                }
                console.log(`▸ Downloading Core ML ${relativePath}...`);
                await downloadAsset({ assetSrc: `registry://${info.registrySource}/${info.registryPath}` });
            }
            const cached = await getModelInfo({ name });
            const sourcePath = cached.cacheFiles.find((file) => file.isCached)?.path;
            if (!sourcePath)
                throw new Error(`Core ML file is missing from the SDK cache: ${name}`);
            const targetPath = join(bundleDir, relativePath);
            await mkdir(join(targetPath, '..'), { recursive: true });
            await link(sourcePath, targetPath);
        }
        return { modelSrc, dir };
    }
    catch (error) {
        await rm(dir, { recursive: true, force: true });
        throw error;
    }
}
async function main() {
    if (!modelPath) {
        throw new Error('Usage: bun run examples/asr/parakeet-unified-coreml.ts ' +
            '<regenerated-unified-gguf> [16-kHz-wav-file]');
    }
    let staged;
    let modelId;
    try {
        staged = await stageModel(modelPath);
        console.log('▸ Loading Parakeet Unified with its Core ML encoder sidecar...');
        modelId = await loadModel({
            modelSrc: staged.modelSrc,
            modelType: 'parakeet-transcription',
            modelConfig: { useGPU: true }
        });
        console.log('▸ Transcribing audio...');
        const text = await transcribe({ modelId, audioChunk: audioPath });
        if (!text.trim())
            throw new Error('Parakeet Unified returned an empty transcription');
        console.log(text);
    }
    finally {
        try {
            if (modelId)
                await unloadModel({ modelId });
        }
        finally {
            if (staged)
                await rm(staged.dir, { recursive: true, force: true });
        }
    }
}
main().catch((error) => {
    console.error('✖', error);
    process.exitCode = 1;
});

Parakeet Indic Conformer

The following script shows an example of Indic-language transcription using the Parakeet Indic Conformer CTC model from a WAV file:

parakeet-indic-conformer-filesystem.js
/**
 * Indic Conformer CTC transcription from a WAV file.
 *
 * Usage:
 *   bun run examples/asr/parakeet-indic-conformer-filesystem.ts <wav-file> <language> [gguf]
 *
 * Loads a single GGUF checkpoint (`PARAKEET_INDIC_CONFORMER_600M_Q8_0` by
 * default) and transcribes with the batch `transcribe` API. `language` is
 * required (e.g. `hi`, `ta`) because Indic Conformer CTC masks the vocab
 * with `parakeet.ctc.lang_*` ranges. English Parakeet CTC ignores this field.
 *
 * Audio should be 16 kHz mono PCM in a WAV container.
 */
import { loadModel, unloadModel, transcribe, PARAKEET_INDIC_CONFORMER_600M_Q8_0 } from '@qvac/sdk';
const args = process.argv.slice(2);
if (!args[0] || !args[1]) {
    console.error('Usage: bun run examples/asr/parakeet-indic-conformer-filesystem.ts ' +
        '<wav-file> <language> [indic-conformer-gguf]');
    console.error('\nExample: ... filesystem.ts speech.wav hi');
    console.error('If the model path is omitted, defaults to the registry model.');
    process.exit(1);
}
const audioFilePath = args[0];
const language = args[1];
const parakeetModelSrc = args[2] ?? PARAKEET_INDIC_CONFORMER_600M_Q8_0;
try {
    console.log('▸ Loading Indic Conformer CTC model...');
    const modelId = await loadModel({
        modelSrc: parakeetModelSrc,
        modelType: 'parakeet-transcription',
        modelConfig: { language },
        onProgress: (p) => {
            const mb = (n) => (n / 1e6).toFixed(1);
            const line = `▸ Downloading ${p.percentage.toFixed(0)}% (${mb(p.downloaded)}/${mb(p.total)} MB)`;
            process.stderr.write(process.stderr.isTTY ? `\r${line}` : `${line}\n`);
            if (p.percentage >= 100)
                process.stderr.write('\n');
        }
    });
    console.log(`▸ Indic Conformer CTC model loaded with ID: ${modelId}`);
    console.log(`▸ Language: ${language}`);
    console.log('▸ Transcribing audio...');
    const text = await transcribe({ modelId, audioChunk: audioFilePath });
    console.log(text);
    console.log('▸ Unloading model...');
    await unloadModel({ modelId });
    console.log('▸ Done');
}
catch (error) {
    console.error('✖', error);
    process.exit(1);
}

Parakeet Sortformer

The following script shows an example of speaker diarization using the Parakeet Sortformer model, followed by per-segment transcription with the TDT model:

parakeet-sortformer.js
/**
 * Parakeet Sortformer diarization + TDT transcription pipeline.
 *
 * Usage:
 *   bun run examples/asr/parakeet-sortformer.ts [sortformer-gguf] [wav-file]
 *
 * Two-step flow: Sortformer v2.1 diarizes the audio, then TDT transcribes each
 * speaker segment. Defaults to registry GGUFs and
 * `examples/audio/diarization-sample-16k.wav`. For live streaming + AOSC, see
 * `parakeet-sortformer-streaming.ts`.
 *
 * Sample audio is in the QVAC source repo but not the published npm package.
 * Download the default file into `examples/audio/`:
 *   https://github.com/tetherto/qvac/blob/main/packages/sdk/examples/audio/diarization-sample-16k.wav
 */
import { loadModel, unloadModel, transcribe, PARAKEET_TDT_0_6B_V3_Q8_0, PARAKEET_SORTFORMER_4SPK_V2_1_Q8_0 } from '@qvac/sdk';
import { dirname, join } from 'path';
import { fileURLToPath } from 'url';
import { readFileSync, writeFileSync, mkdirSync } from 'fs';
import { tmpdir } from 'os';
const __dirname = dirname(fileURLToPath(import.meta.url));
const args = process.argv.slice(2);
const sortformerSrc = args[0] ?? PARAKEET_SORTFORMER_4SPK_V2_1_Q8_0;
const defaultAudioPath = join(__dirname, '..', 'audio', 'diarization-sample-16k.wav');
const audioFilePath = args[1] ?? defaultAudioPath;
try {
    // ── Step 1: Diarize with Sortformer ──
    const sfModelId = await loadModel({
        modelSrc: sortformerSrc,
        modelType: 'parakeet-transcription'
    });
    const diarization = await transcribe({
        modelId: sfModelId,
        audioChunk: audioFilePath
    });
    await unloadModel({ modelId: sfModelId });
    const segments = parseDiarization(diarization);
    // ── Step 2: Transcribe each segment with TDT ──
    const tdtModelId = await loadModel({
        modelSrc: PARAKEET_TDT_0_6B_V3_Q8_0
    });
    const pcm = readPcm(audioFilePath);
    const sliceDir = join(tmpdir(), `qvac-diarize-${Date.now()}`);
    mkdirSync(sliceDir, { recursive: true });
    const results = [];
    for (let i = 0; i < segments.length; i++) {
        const seg = segments[i];
        const slicePath = join(sliceDir, `seg-${i}.wav`);
        if (!writeWavSlice(pcm, seg.start, seg.end, slicePath)) {
            results.push({ ...seg, text: '[No speech detected]' });
            continue;
        }
        const text = await transcribe({
            modelId: tdtModelId,
            audioChunk: slicePath
        });
        results.push({ ...seg, text: text.trim() || '[No speech detected]' });
    }
    await unloadModel({ modelId: tdtModelId });
    // ── Step 3: Merge consecutive same-speaker segments and print ──
    const merged = mergeSpeakers(results);
    console.log('\n▸ Diarized transcription');
    for (const entry of merged) {
        console.log(`Speaker ${entry.speaker} (${entry.start.toFixed(2)}s - ${entry.end.toFixed(2)}s):`);
        console.log(`  ${entry.text}\n`);
    }
    console.log('▸ Done');
}
catch (error) {
    console.error('✖', error);
    process.exit(1);
}
// ── Helpers ──
function parseDiarization(text) {
    const segs = [];
    for (const line of text.split('\n')) {
        const m = line.match(/Speaker (\d+): ([\d.]+)s - ([\d.]+)s/);
        if (m)
            segs.push({ speaker: +m[1], start: +m[2], end: +m[3] });
    }
    return segs.sort((a, b) => a.start - b.start);
}
function readPcm(wavPath) {
    const buf = readFileSync(wavPath);
    const dataOffset = buf.indexOf('data') + 4;
    return buf.subarray(dataOffset + 4, dataOffset + 4 + buf.readUInt32LE(dataOffset));
}
function writeWavSlice(pcm, startSec, endSec, outPath) {
    const SR = 16000;
    const BPS = 2;
    const startByte = Math.floor(startSec * SR) * BPS;
    const endByte = Math.min(Math.ceil(endSec * SR) * BPS, pcm.length);
    if (startByte >= endByte)
        return false;
    const slice = pcm.subarray(startByte, endByte);
    const hdr = Buffer.alloc(44);
    hdr.write('RIFF', 0);
    hdr.writeUInt32LE(36 + slice.length, 4);
    hdr.write('WAVEfmt ', 8);
    hdr.writeUInt32LE(16, 16);
    hdr.writeUInt16LE(1, 20);
    hdr.writeUInt16LE(1, 22);
    hdr.writeUInt32LE(SR, 24);
    hdr.writeUInt32LE(SR * BPS, 28);
    hdr.writeUInt16LE(BPS, 32);
    hdr.writeUInt16LE(16, 34);
    hdr.write('data', 36);
    hdr.writeUInt32LE(slice.length, 40);
    writeFileSync(outPath, Buffer.concat([hdr, slice]));
    return true;
}
function mergeSpeakers(entries) {
    const out = [];
    for (const e of entries) {
        const last = out[out.length - 1];
        if (last && last.speaker === e.speaker) {
            last.text += ' ' + e.text;
            last.end = e.end;
        }
        else {
            out.push({ ...e });
        }
    }
    return out;
}

Tip: all examples throughout this documentation are self-contained and runnable. For instructions on how to run them, see the JS/TS quickstart or the Python quickstart.

On this page

Ask anything about QVAC.