New: TranslatePsy-AfriSLM translates directly between 19 African languages, offline.
QVAC Logo
SDKRPC servers
v0.21, the current release

RPC servers

Start and discover native RPC servers, then use their devices for model inference.

Overview

Use startRpcServer() to serve native devices from a QVAC worker and discoverRpcServers() to find endpoints on a shared discovery topic. Load a model with the selected endpoints to run inference across their devices.

Use RPC endpoints only on a trusted private network. Native RPC traffic is unencrypted and unauthenticated. A discovery topic does not authenticate peers.

Enable serving

Serving requires @qvac/ggml-rpc-server and a worker rebuilt with this setting in qvac.config.json:

{
  "rpcServerProvider": "@qvac/sdk/ggml-rpc-server/provider"
}

Rebuild the worker with qvac bundle sdk. See Configuration for the config file format.

Python callers must select the generated worker with QVAC_WORKER_PATH or Client(worker_path=...). See Python SDK RPC server setup for installation, bundling, and run commands.

Start and stop a server

startRpcServer() defaults to 127.0.0.1 and allocates a free TCP port when port is omitted. It returns serverId, url, and rdmaCapable.

To serve another machine, set host to your private IPv4 address and explicitly set allowNonLoopbackHost: true. Set discoveryTopic to advertise the ready endpoint under a shared topic.

Unload models using the endpoint before calling stopRpcServer({ serverId }). This withdraws the announcement and waits for the native server to stop. Native stop under Bare has no deadline.

Discover and select devices

discoverRpcServers({ topic, timeoutMs }) returns candidates with a url, rdmaAvailable, and devices. Each device has an endpoint-local index, freeMemory, and totalMemory, with memory values in bytes. The search budget defaults to 5000 milliseconds, includes TCP probes, and is capped at 30000 milliseconds.

rdmaAvailable describes the server's advertised RDMA availability. Actual transport depends on client support and negotiation.

Choose and order the endpoints before calling getRpcDeviceMap(selected). Pass those same endpoints, in the same order, as the comma-separated modelConfig["rpc-servers"] value to loadModel(). The map assigns RPC0, RPC1, and subsequent aliases across every device in endpoint order. Aliases start at RPC0 for each model load, including when you reuse a worker.

Use the chosen aliases in modelConfig.devices. Set "split-mode": "layer" to split layers across devices, and order "tensor-split" weights to match the selected devices. See Text generation for the inference API.

Cancellation

Pending server starts and discovery calls expose a synchronous requestId on their returned promise. Pass it to cancel() to cancel that call. See Cancellation. After startup completes, stop the server with stopRpcServer().

Examples

Serve and discover endpoints

The following script starts an RPC server on a private IPv4 address and shared topic, then waits for you to stop it after clients unload their models:

rpc-server.js
// Rebuild the worker with
// rpcServerProvider: '@qvac/sdk/ggml-rpc-server/provider' in qvac.config.json.
// npx tsx rpc-server.ts 10.0.0.2 my-private-rpc-group
import { createInterface } from 'node:readline/promises';
import { startRpcServer, stopRpcServer, close } from '@qvac/sdk';
try {
    const host = process.argv[2];
    const topic = process.argv[3];
    if (!host || !topic)
        throw new Error('Usage: rpc-server.ts <private-ipv4> <topic>');
    const server = await startRpcServer({
        host,
        allowNonLoopbackHost: true,
        discoveryTopic: topic
    });
    const input = createInterface({ input: process.stdin, output: process.stderr });
    try {
        console.error(`▸ Serving ${server.url}; RDMA capable: ${server.rdmaCapable}`);
        await input.question('▸ Press Enter to stop after clients unload their models.\n');
    }
    finally {
        input.close();
        await stopRpcServer({ serverId: server.serverId });
    }
}
catch (error) {
    console.error('✖', error);
    process.exitCode = 1;
}
finally {
    await close();
}

The Python script has separate serve and discover modes. See Python SDK for its worker requirement and typed methods.

Inference across two servers

The following script discovers two idle RPC servers, maps their device aliases, and streams a completion with layer splitting and equal device weights:

rpc-inference.js
// npx tsx rpc-inference.ts my-private-rpc-group
import { discoverRpcServers, getRpcDeviceMap, loadModel, completion, unloadModel, close, LLAMA_3_2_1B_INST_Q4_0 } from '@qvac/sdk';
try {
    const topic = process.argv[2];
    if (!topic)
        throw new Error('Usage: rpc-inference.ts <topic>');
    // Prefer server RDMA offers before assigning aliases. Transport is negotiated at load time.
    const selected = (await discoverRpcServers({ topic, timeoutMs: 5000 }))
        .sort((a, b) => Number(b.rdmaAvailable) - Number(a.rdmaAvailable) || a.url.localeCompare(b.url))
        .slice(0, 2);
    if (selected.length !== 2)
        throw new Error('Two idle RPC servers are required');
    const devices = getRpcDeviceMap(selected);
    if (!devices.length)
        throw new Error('The selected servers expose no devices');
    console.error(`▸ Using ${devices.length} devices across ${selected.length} servers`);
    const modelId = await loadModel({
        modelSrc: LLAMA_3_2_1B_INST_Q4_0,
        modelConfig: {
            device: 'gpu',
            'rpc-servers': selected.map((server) => server.url).join(','),
            devices: devices.map((device) => device.alias).join(','),
            'split-mode': 'layer',
            'tensor-split': devices.map(() => '1').join(',')
        },
        onProgress: (progress) => {
            process.stderr.write(`\r▸ Downloading ${progress.percentage.toFixed(0)}%`);
        }
    });
    try {
        process.stderr.write('\n');
        const run = completion({
            modelId,
            history: [{ role: 'user', content: 'Explain distributed inference in one sentence.' }],
            stream: true
        });
        for await (const token of run.tokenStream)
            process.stdout.write(token);
        process.stdout.write('\n');
    }
    finally {
        await unloadModel({ modelId });
    }
}
catch (error) {
    console.error('✖', error);
    process.exitCode = 1;
}
finally {
    await close();
}

All examples are self-contained and runnable. See the JS/TS quickstart or Python quickstart for instructions.

On this page

Ask anything about QVAC.