New: VisionPsy-Nano, a 460M vision model that outperforms models twice its size.
QVAC Logo

Plugin system

Enable and disable built-in AI capabilities, and add new ones via custom plugins.

Overview

Each QVAC AI capability maps to a built-in plugin in the SDK. This lets you enable only what you need for your project and reduce your application's final bundle size.

In addition, you can add custom plugins — both your own and community ones — that extend QVAC's capabilities. In both cases, you'll use the QVAC configuration file and then bundle either with QVAC CLI or programmatically via @qvac/sdk/commands.

Built-in plugins

Catalog

All the built-in plugins you can select, along with the AI tasks that depend on each one:

PluginUse in qvac.config.*AI tasks that require it
LLM@qvac/sdk/llamacpp-completion/pluginText generation; multimodal; RAG ; Fine-tuning
Embeddings@qvac/sdk/llamacpp-embedding/pluginText embeddings; RAG
ASR with customized Whisper engine@qvac/sdk/whispercpp-transcription/pluginTranscription
ASR with Parakeet@qvac/sdk/parakeet-transcription/pluginTranscription
BCI transcription@qvac/sdk/bci-whispercpp-transcription/pluginTranscription
BCI transcription@qvac/sdk/bci-whispercpp-transcription/pluginTranscription
NMT@qvac/sdk/nmtcpp-translation/pluginTranslation
TTS@qvac/sdk/onnx-tts/pluginText-to-Speech
OCR@qvac/sdk/ggml-ocr/pluginOCR
Diffusion@qvac/sdk/sdcpp-generation/pluginImage generation
AudioGen@qvac/sdk/audiogen-ggml/pluginMusic generation

Enabling

In your qvac.config.*, add the built-in plugins you’ll need in your project. For example:

qvac.config.json
{
  "plugins": [
    "@qvac/sdk/llamacpp-completion/plugin",
    "@qvac/sdk/ggml-ocr/plugin"
  ]
}

When developing for desktop environment, use the QVAC CLI to (re)bundle the SDK only with the selected plugins:

qvac bundle sdk

Only the selected plugins are included in your bundle, significantly reducing your application size. If plugins is omitted or an empty array, it bundles all built-in plugins by default.

Tips:

Custom plugins

Custom plugins are consumed as npm packages. Install the package, enable it in your qvac.config.* file like any built-in plugin, then import and use its API alongside the SDK's regular API.

Enabling

Install the plugin package in your project. For example:

npm i qvac-echo-plugin

In your qvac.config.*, add its <package_name>/plugin specifier alongside any built-in plugins. For example:

qvac.config.json
{
  "plugins": [
    "@qvac/sdk/llamacpp-completion/plugin",
    "qvac-echo-plugin/plugin"
  ]
}

When developing for desktop environment, use the SDK CLI to (re)bundle the SDK with all listed plugins:

qvac bundle sdk

Tip: when adding one or more custom plugins, you must also add all the built-in plugins you will need to use.

Usage

Custom plugins are consumed as npm packages. Install the package, enable it in your qvac.config.* file like any built-in plugin, then import and use its API alongside the SDK’s regular API.

import { loadModel, unloadModel } from "@qvac/sdk";

// 1. Import the API you need from the custom plugin package:
import { echo, echoStream } from "qvac-echo-plugin";

// 2. Load the model(s) required by the custom plugin:
const modelId = await loadModel({
  modelSrc: "/path/to/echo-model.bin",
  modelType: "echo",
});

// 3. Call functions exposed by the custom plugin API:
const result = await echo({ modelId, message: "Hello, plugin system!" });
console.log(result);

for await (const char of echoStream({ modelId, message: "Streaming test!" })) {
  process.stdout.write(char);
}

await unloadModel({ modelId });

The Python SDK exposes plugin invocation through the same worker contract:

plugins.py
"""Python port of packages/sdk/examples/plugins.ts (plugin invocation).

The low-level plugin API. A model's `modelType` selects a plugin; each plugin
exposes named `handler`s with their own request/response schemas.
`invoke_plugin(t, model_id, handler, params)` runs one handler and returns its
result; `invoke_plugin_stream(...)` yields streamed chunks.

This example uses the `custom-echo-plugin` fixture (modelType `echo`), whose
handlers are `echo` ({message} -> {message}) and `echoStream` ({message} ->
{chunk} chunks). To run it you need a worker bundled with that plugin; point
the loader at it via argv, or adapt the handler/params to your own plugin.
(A production model that exposes handlers is VLA — see vla.py, which drives the
`vlaRun` / `vlaHparams` handlers through typed wrappers.)

RUN: python examples/plugins.py [model-src]
"""

from __future__ import annotations

import asyncio
import sys

from _common import print_progress

from tetherto.qvac_sdk import (
    Client,
    invoke_plugin,
    invoke_plugin_stream,
    load_model,
    unload_model,
)


async def main() -> int:
    if len(sys.argv) < 2:
        print(
            "Usage: python examples/plugins.py <model-src>\n"
            "  <model-src> must resolve to a model whose plugin exposes the "
            "'echo'/'echoStream' handlers (the custom-echo-plugin fixture).",
            file=sys.stderr,
        )
        return 1
    model_src = sys.argv[1]

    async with Client() as client:
        t = client.transport
        try:
            print("▸ Loading plugin-backed model...")
            model_id = await load_model(
                t,
                model_src=model_src,
                model_type="echo",
                on_progress=print_progress,
            )
            print(f"▸ Model loaded: {model_id}")

            print("\n▸ 1. invoke_plugin (request/reply handler 'echo')")
            result = await invoke_plugin(
                t, model_id, "echo", params={"message": "hello from the python sdk"}
            )
            print(f"▸ echo -> {result}")

            print("\n▸ 2. invoke_plugin_stream (streaming handler 'echoStream')")
            print("▸ chunks: ", end="")
            async for chunk in invoke_plugin_stream(
                t, model_id, "echoStream", params={"message": "streamed message"}
            ):
                sys.stdout.write(str(getattr(chunk, "chunk", chunk)))
                sys.stdout.flush()
            print()

            await unload_model(t, model_id)
            print("▸ Model unloaded")
        except Exception as error:
            print(f"✖ {error}", file=sys.stderr)
            return 1
    return 0


if __name__ == "__main__":
    sys.exit(asyncio.run(main()))

Runtime registration on Bare

The plugins array in qvac.config.* is bundle-time configuration — it controls which addons get packed into your worker. That is separate from runtime registration, which determines the plugins live in the worker process when an SDK call runs.

On Node.js and Expo the SDK spawns a worker that auto-registers the full built-in set, so you never register manually. Bare runs in-process with no spawned worker, so nothing auto-registers — register the plugins you use before the first SDK call:

import { plugins } from "@qvac/bare-sdk";
import { llmPlugin } from "@qvac/bare-sdk/llamacpp-completion/plugin";

const sdk = plugins([llmPlugin]); // or registerPlugin(llmPlugin) from "@qvac/bare-sdk/plugins"

Calls made before any plugin is registered raise WorkerPluginsNotRegisteredError.

For direct Bare usage we recommend @qvac/bare-sdk — the slim distribution built for explicit assembly. @qvac/sdk also runs on Bare; import the same plugins from @qvac/sdk/<capability>/plugin instead.

Notes

  • In qvac.config.*, if plugins is omitted or set to an empty array, the SDK bundles all built-in plugins. If plugins is set, it bundles only the listed plugins.
  • Each plugin maps to one QVAC addon.
  • Write a custom plugin.
  • Programmatic bundling: you can also prepare and validate the SDK bundle programmatically via API from a Node script. For example:
import { bundleSdk, verifyBundle } from "@qvac/sdk/commands";

await bundleSdk({
  projectRoot: process.cwd(),
  configPath: "./qvac.config.json",
  quiet: true,
});

const result = await verifyBundle({
  projectRoot: process.cwd(),
  addonsSource: "./qvac/worker.bundle.js",
  hosts: ["android-arm64", "ios-arm64"],
});

On this page

Ask anything about QVAC.