New: TranslatePsy-AfriSLM translates directly between 19 African languages, offline.
QVAC Logo
SDKRelease notes
v0.21, the current release

SDK Release Notes — v0.21.x (latest)

Release notes for QVAC SDK v0.21.0.

v0.21.0

@qvac/sdk

📦 NPM: https://www.npmjs.com/package/@qvac/sdk/v/0.21.0

QVAC SDK 0.21.0 sits on @qvac/inference@0.21.0. A model can split across machines through llama.cpp's RPC backend, Ternary Bonsai 2 27B runs in QVAC's engine, tools can defer their schemas behind tool_search, and TTS/ASR pick up MOSS plus Parakeet Core ML on Apple. Mobile bundles can auto-install addon platform packages, and verifyBundle checks engines.bare against the Bare runtime each host actually runs. Catalog constant names for BitNet, Llama tool-calling, and Indic Parakeet change. assessModelFit now takes loadModel fields instead of a separate workload object, infers modelType from modelSrc when omitted, and the load-time fit probe runs on every engine that ships a fitter.

The committed range is ^0.21.0. tetherto-qvac-sdk publishes at the same version.

Breaking Changes

assessModelFit uses loadModel parameters

Candidates are described the same way as a load: modelSrc, modelType, and modelConfig. The old model / workload / artifacts object is rejected.

Before:

await assessModelFit({
  models: [
    {
      model: LLAMA_3_2_1B,
      workload: { kind: 'llm', contextTokens: 4096 },
      artifacts: [MMPROJ_F16]
    }
  ]
})

After:

await assessModelFit({
  models: [
    {
      modelSrc: LLAMA_3_2_1B,
      modelType: 'llamacpp-completion',
      modelConfig: { ctx_size: 4096, projectionModelSrc: MMPROJ_F16 }
    }
  ]
})

modelType is optional and inferred from modelSrc when omitted, the same way loadModel infers it. Required only where the source names no engine. A load whose sources are all config fields may omit modelSrc. Audio workload.windowMs becomes modelConfig.duration_ms. workload.batch has no equivalent and is gone.

Flattened fit-probe projection

getLoadedModelInfo().fitProbe.projection is a flat byte breakdown. projection.devices and NativeProbeDevice are gone. Totals are summed across devices; deviceName names the first.

Before:

info.fitProbe?.projection?.devices

After:

info.fitProbe?.projection?.deviceBytes
info.fitProbe?.projection?.hostBytes
info.fitProbe?.projection?.weightsBytes
info.fitProbe?.projection?.contextBytes
info.fitProbe?.projection?.computeBytes
info.fitProbe?.projection?.deviceName
Catalog constant names

BitNet instructed TQ2_0 exports are retagged as base. Llama tool-calling 1B moves from Q4_K to Q4_K_M. Indic Parakeet Conformer CTC constants are replaced by the 600M GGUFs.

Before:

import { BITNET_0_7B_INST_TQ2_0 } from '@qvac/sdk'
import { LLAMA_TOOL_CALLING_1B_INST_Q4_K } from '@qvac/sdk'
import { PARAKEET_INDIC_CONFORMER_CTC_Q4_0 } from '@qvac/sdk'

After:

import { BITNET_0_7B_BASE_TQ2_0 } from '@qvac/sdk'
import { LLAMA_TOOL_CALLING_1B_INST_Q4_K_M } from '@qvac/sdk'
import { PARAKEET_INDIC_CONFORMER_600M_Q4_0 } from '@qvac/sdk'

New APIs

Clustered inference over RPC

startRpcServer starts a worker-owned llama.cpp RPC server. discoverRpcServers finds peers on a private topic and getRpcDeviceMap builds the devices string loadModel expects. One model splits across several machines, phones included, by layer or within a layer. Only the initiator needs the model file.

import {
  startRpcServer,
  stopRpcServer,
  discoverRpcServers,
  getRpcDeviceMap,
  loadModel,
  unloadModel
} from '@qvac/sdk'

const server = await startRpcServer({
  host: '10.0.0.2',
  allowNonLoopbackHost: true,
  discoveryTopic: 'my-private-rpc-group'
})

const selected = (await discoverRpcServers({ topic: 'my-private-rpc-group' }))
  .sort((a, b) => a.url.localeCompare(b.url))
  .slice(0, 2)
const devices = getRpcDeviceMap(selected)
const modelId = await loadModel({
  modelSrc,
  modelConfig: {
    device: 'gpu',
    'rpc-servers': selected.map((s) => s.url).join(','),
    devices: devices.map((d) => d.alias).join(','),
    'split-mode': 'layer',
    'tensor-split': devices.map(() => '1').join(',')
  }
})
await unloadModel({ modelId })
await stopRpcServer({ serverId: server.serverId })

Workers bind to loopback unless allowNonLoopbackHost is set. The channel is unauthenticated. Prebuilds use TCP; RDMA needs client and worker rebuilt on Linux with RDMA enabled. The current limit is 16 devices.

Deferred tool loading

Set deferLoading: true on a tool (or on an MCP client entry) so the first prompt carries only always-loaded tools, a built-in tool_search, and a short catalog of names and one-line descriptions grouped by group. The model searches once, the SDK appends the matching definitions as a tool result, and the model calls them natively on its next step. Tools without the flag work as they do today. completion() stays stateless and reads loaded tools back from history. qvac serve maps defer_loading off the OpenAI tool JSON and runs the same loop.

completion({
  modelId,
  history,
  tools: [
    { type: 'function', name: 'get_weather', description: '...', parameters },
    {
      type: 'function',
      name: 'create_issue',
      description: 'Open a new issue on a repository',
      group: 'github',
      deferLoading: true,
      parameters
    }
  ]
})

tool_search is reserved. generationParams.tool_choice cannot name a deferred tool.

MOSS TTS

tts-ggml 0.10.0 adds the MOSS engine. MOSS-TTS-v1.5 is directable (explicit pause and duration, Pinyin/IPA, optional voice cloning). MOSS-TTSD synthesizes 1 to 5 speakers in one pass. The Delay engine streams chunks so playback can start while generation continues.

import {
  loadModel,
  textToSpeech,
  TTS_DELAY_LLM_MOSS_TTS_F16,
  TTS_CODEC_DECODER_MOSS_TTS_F16,
  TTS_CODEC_ENCODER_MOSS_TTS_F16
} from '@qvac/sdk'

const modelId = await loadModel({
  modelSrc: TTS_DELAY_LLM_MOSS_TTS_F16,
  modelType: 'tts',
  modelConfig: {
    ttsEngine: 'moss',
    mossCodecDecoderModelSrc: TTS_CODEC_DECODER_MOSS_TTS_F16,
    mossCodecEncoderModelSrc: TTS_CODEC_ENCODER_MOSS_TTS_F16,
    referenceAudioSrc: '/voices/speaker-24k.wav',
    streamChunkTokens: 25
  }
})

On Apple, Audio8 now stages its Core ML codec sidecar and reports codecSidecarLoaded / codecOnCoreml on synthesis stats.

ABot-World layer streaming

World-mode modelConfig.world accepts paramsBackend, maxVram, streamLayers, and the rest of the placement fields so an ABot-World graph can stream layers the same way other diffusion loads do.

Mobile host prebuilds and Bare runtime checks

ensureHostPrebuilds and bundleSdk({ installMissingPrebuilds: true }) add the Android/iOS platform packages the split speech addons need, pinned to each addon's version. qvac bundle sdk installs by default; --no-install skips it.

verifyBundle and qvac doctor compare each host's actual Bare runtime (from react-native-bare-kit on mobile) against engines.bare, and print the kit release or override to apply.

Loaded-model fit probe

getLoadedModelInfo can include fitProbe: verdict, a flat projection (deviceBytes, hostBytes, and the breakdown fields the engine filled), and the placement the fitter resolved. The probe runs on every engine that ships a fitter. assessModelFit model results can carry device and reasons. fitStubBudgetMs on config (default 40000) budgets the registry fetch of the weightless description.

BCI stream placement and diagnostics

bciTranscribe terminal frames can carry diagnostics. Streamed BCI segments expose windowStartTimestep so window-local timestamps can be placed on the stream timeline.

Features

Ternary Bonsai 2 27B (PTQ1_0 at 5.95 GB, PQ2_0 at 7.21 GB) loads through the qvac-fabric engine with the matching rotation applied at run time.

Parakeet GGUFs on Apple download their Core ML sidecars beside the weights. ACE-Step AudioGen packings and additional CosyVoice3 flow/hift/LLM constants are on the catalog.

Llama gpu_layers is unset by default so qvac-fabric can place layers to free device memory. Setting gpu_layers still pins the count and aborts that fit.

Hyperdrive model downloads stay inside the model cache. NMT translate() stats are per-request. Fused TTS export names come from registry tags. bare-runtime is ^1.30.3 so a retained lockfile cannot keep 1.24.x next to engines.bare >=1.30.3.

Bug Fixes

Loaded deferred tools run under the tool-call grammar instead of text parsing on the search-to-call step.

Model Changes

Added
AUDIOGEN_ACESTEP_5HZ_LM_0_6B_BF16
AUDIOGEN_ACESTEP_V15_BASE_Q4_K_M
BITNET_0_7B_BASE_TQ2_0
BITNET_1B_BASE_TQ2_0
BITNET_B1_58_3B_BASE_TQ2_0
LLAMA_TOOL_CALLING_1B_INST_Q4_K_M
PARAKEET_INDIC_CONFORMER_600M_F16
PARAKEET_INDIC_CONFORMER_600M_Q4_0
PARAKEET_INDIC_CONFORMER_600M_Q8_0
PARAKEET_TDT_1_1B_F16
PARAKEET_TDT_1_1B_Q8_0
TERNARY_BONSAI_2_27B_MULTIMODAL_PQ2_0
TERNARY_BONSAI_2_27B_MULTIMODAL_PTQ1_0
TTS_CODEC_DECODER_MOSS_TTS_F16
TTS_CODEC_ENCODER_MOSS_TTS_F16
TTS_COSYVOICE3_FLOW_COSYVOICE_BF16
TTS_COSYVOICE3_FLOW_COSYVOICE_FP16
TTS_COSYVOICE3_FLOW_COSYVOICE_Q4_0
TTS_COSYVOICE3_FLOW_COSYVOICE_Q8_0
TTS_COSYVOICE3_HIFT_COSYVOICE_FP16
TTS_COSYVOICE3_LLM_COSYVOICE_FUSED_Q8_0
TTS_COSYVOICE3_LLM_COSYVOICE_Q4_0
TTS_DELAY_LLM_MOSS_TTS_F16
Removed
BITNET_0_7B_INST_TQ2_0
BITNET_1B_INST_TQ2_0
BITNET_B1_58_3B_INST_TQ2_0
LLAMA_TOOL_CALLING_1B_INST_Q4_K
PARAKEET_INDIC_CONFORMER_CTC_F16
PARAKEET_INDIC_CONFORMER_CTC_Q4_0
PARAKEET_INDIC_CONFORMER_CTC_Q8_0

On this page

Ask anything about QVAC.