---
title: "RPC servers"
canonical: https://docs.qvac.tether.io/sdk/p2p-capabilities/rpc-servers/
collection: "SDK"
package: "@qvac/sdk"
line: v0.21
current_line: true
---

# RPC servers (/sdk/p2p-capabilities/rpc-servers)



## Overview

Use `startRpcServer()` to serve native devices from a QVAC worker and `discoverRpcServers()` to find endpoints on a shared discovery topic. Load a model with the selected endpoints to run inference across their devices.

<Callout type="warn">
  Use RPC endpoints only on a trusted private network. Native RPC traffic is unencrypted and unauthenticated. A discovery topic does not authenticate peers.
</Callout>

## Enable serving

Serving requires `@qvac/ggml-rpc-server` and a worker rebuilt with this setting in `qvac.config.json`:

```json
{
  "rpcServerProvider": "@qvac/sdk/ggml-rpc-server/provider"
}
```

Rebuild the worker with `qvac bundle sdk`. See [Configuration](/sdk/configuration) for the config file format.

Python callers must select the generated worker with `QVAC_WORKER_PATH` or `Client(worker_path=...)`. See [Python SDK RPC server setup](/sdk/python-sdk#rpc-servers) for installation, bundling, and run commands.

## Start and stop a server

`startRpcServer()` defaults to `127.0.0.1` and allocates a free TCP port when `port` is omitted. It returns `serverId`, `url`, and `rdmaCapable`.

To serve another machine, set `host` to your private IPv4 address and explicitly set `allowNonLoopbackHost: true`. Set `discoveryTopic` to advertise the ready endpoint under a shared topic.

Unload models using the endpoint before calling `stopRpcServer({ serverId })`. This withdraws the announcement and waits for the native server to stop. Native stop under Bare has no deadline.

## Discover and select devices

`discoverRpcServers({ topic, timeoutMs })` returns candidates with a `url`, `rdmaAvailable`, and `devices`. Each device has an endpoint-local `index`, `freeMemory`, and `totalMemory`, with memory values in bytes. The search budget defaults to 5000 milliseconds, includes TCP probes, and is capped at 30000 milliseconds.

`rdmaAvailable` describes the server's advertised RDMA availability. Actual transport depends on client support and negotiation.

Choose and order the endpoints before calling `getRpcDeviceMap(selected)`. Pass those same endpoints, in the same order, as the comma-separated `modelConfig["rpc-servers"]` value to [`loadModel()`](/sdk/reference/api#loadmodel). The map assigns `RPC0`, `RPC1`, and subsequent aliases across every device in endpoint order. Aliases start at `RPC0` for each model load, including when you reuse a worker.

Use the chosen aliases in `modelConfig.devices`. Set `"split-mode": "layer"` to split layers across devices, and order `"tensor-split"` weights to match the selected devices. See [Text generation](/sdk/ai-capabilities/text-generation) for the inference API.

## Cancellation

Pending server starts and discovery calls expose a synchronous `requestId` on their returned promise. Pass it to [`cancel()`](/sdk/reference/api#cancel) to cancel that call. See [Cancellation](/sdk/runtime/cancellation). After startup completes, stop the server with `stopRpcServer()`.

## Examples

### Serve and discover endpoints

The following script starts an RPC server on a private IPv4 address and shared topic, then waits for you to stop it after clients unload their models:

<Tabs>
  <Tab value="js" label="JavaScript" default>
    <WrapCode>
      ```js file=<rootDir>/packages/sdk/dist/examples/rpc-server.js title="rpc-server.js" lineNumbers
      // Rebuild the worker with
      // rpcServerProvider: '@qvac/sdk/ggml-rpc-server/provider' in qvac.config.json.
      // npx tsx rpc-server.ts 10.0.0.2 my-private-rpc-group
      import { createInterface } from 'node:readline/promises';
      import { startRpcServer, stopRpcServer, close } from '@qvac/sdk';
      try {
          const host = process.argv[2];
          const topic = process.argv[3];
          if (!host || !topic)
              throw new Error('Usage: rpc-server.ts <private-ipv4> <topic>');
          const server = await startRpcServer({
              host,
              allowNonLoopbackHost: true,
              discoveryTopic: topic
          });
          const input = createInterface({ input: process.stdin, output: process.stderr });
          try {
              console.error(`▸ Serving ${server.url}; RDMA capable: ${server.rdmaCapable}`);
              await input.question('▸ Press Enter to stop after clients unload their models.\n');
          }
          finally {
              input.close();
              await stopRpcServer({ serverId: server.serverId });
          }
      }
      catch (error) {
          console.error('✖', error);
          process.exitCode = 1;
      }
      finally {
          await close();
      }
      ```
    </WrapCode>
  </Tab>

  <Tab value="ts" label="TypeScript">
    <WrapCode>
      ```ts file=<rootDir>/packages/sdk/examples/rpc-server.ts title="rpc-server.ts" lineNumbers
      // Rebuild the worker with
      // rpcServerProvider: '@qvac/sdk/ggml-rpc-server/provider' in qvac.config.json.
      // npx tsx rpc-server.ts 10.0.0.2 my-private-rpc-group
      import { createInterface } from 'node:readline/promises'
      import { startRpcServer, stopRpcServer, close } from '@qvac/sdk'

      try {
        const host = process.argv[2]
        const topic = process.argv[3]
        if (!host || !topic) throw new Error('Usage: rpc-server.ts <private-ipv4> <topic>')
        const server = await startRpcServer({
          host,
          allowNonLoopbackHost: true,
          discoveryTopic: topic
        })
        const input = createInterface({ input: process.stdin, output: process.stderr })
        try {
          console.error(`▸ Serving ${server.url}; RDMA capable: ${server.rdmaCapable}`)
          await input.question('▸ Press Enter to stop after clients unload their models.\n')
        } finally {
          input.close()
          await stopRpcServer({ serverId: server.serverId })
        }
      } catch (error) {
        console.error('✖', error)
        process.exitCode = 1
      } finally {
        await close()
      }
      ```
    </WrapCode>
  </Tab>

  <Tab value="python" label="Python">
    <WrapCode>
      ```python file=<rootDir>/packages/sdk-python/examples/rpc_servers.py title="rpc_servers.py" lineNumbers
      """Serve or discover RPC endpoints.

      Serving requires @qvac/ggml-rpc-server and a worker rebuilt with
      rpcServerProvider: "@qvac/sdk/ggml-rpc-server/provider" in qvac.config.json.

      python examples/rpc_servers.py serve 10.0.0.2 my-private-rpc-group
      python examples/rpc_servers.py discover my-private-rpc-group
      """

      from __future__ import annotations

      import argparse
      import asyncio
      import sys

      from tetherto.qvac_sdk import Client
      from tetherto.qvac_sdk.methods import (
          discover_rpc_servers,
          start_rpc_server,
          stop_rpc_server,
      )
      from tetherto.qvac_sdk.schemas import (
          DiscoverRpcServersRequest,
          StartRpcServerRequest,
          StopRpcServerRequest,
      )


      async def main() -> int:
          parser = argparse.ArgumentParser(description=__doc__)
          modes = parser.add_subparsers(dest="mode", required=True)
          serve = modes.add_parser("serve")
          serve.add_argument("host")
          serve.add_argument("topic")
          discover = modes.add_parser("discover")
          discover.add_argument("topic")
          args = parser.parse_args()
          try:
              async with Client() as client:
                  transport = client.transport
                  if args.mode == "serve":
                      server = await start_rpc_server(
                          transport,
                          StartRpcServerRequest(
                              type="startRpcServer",
                              host=args.host,
                              allow_non_loopback_host=True,
                              discovery_topic=args.topic,
                          ),
                      )
                      try:
                          print(f"▸ Serving {server.url}", file=sys.stderr)
                          print(
                              "▸ Press Enter after clients unload their models.",
                              file=sys.stderr,
                          )
                          await asyncio.to_thread(input)
                      finally:
                          await stop_rpc_server(
                              transport,
                              StopRpcServerRequest(
                                  type="stopRpcServer", server_id=server.server_id
                              ),
                          )
                  else:
                      result = await discover_rpc_servers(
                          transport,
                          DiscoverRpcServersRequest(
                              type="discoverRpcServers", topic=args.topic, timeout_ms=5000
                          ),
                      )
                      alias = 0
                      # Prefer server RDMA offers before assigning aliases. Transport still needs negotiation.
                      for server in sorted(
                          result.servers, key=lambda item: (not item.rdma_available, item.url)
                      ):
                          print(
                              f"▸ {server.url}; server RDMA available: {server.rdma_available}",
                              file=sys.stderr,
                          )
                          for device in server.devices:
                              print(f"RPC{alias}: {server.url}, device {device.index}")
                              alias += 1
                      if not alias:
                          print("▸ No idle servers found", file=sys.stderr)
          except Exception as error:
              print(f"✖ {error}", file=sys.stderr)
              return 1
          return 0


      if __name__ == "__main__":
          sys.exit(asyncio.run(main()))
      ```
    </WrapCode>
  </Tab>
</Tabs>

The Python script has separate `serve` and `discover` modes. See [Python SDK](/sdk/python-sdk#rpc-servers) for its worker requirement and typed methods.

### Inference across two servers

The following script discovers two idle RPC servers, maps their device aliases, and streams a completion with layer splitting and equal device weights:

<Tabs>
  <Tab value="js" label="JavaScript" default>
    <WrapCode>
      ```js file=<rootDir>/packages/sdk/dist/examples/rpc-inference.js title="rpc-inference.js" lineNumbers
      // npx tsx rpc-inference.ts my-private-rpc-group
      import { discoverRpcServers, getRpcDeviceMap, loadModel, completion, unloadModel, close, LLAMA_3_2_1B_INST_Q4_0 } from '@qvac/sdk';
      try {
          const topic = process.argv[2];
          if (!topic)
              throw new Error('Usage: rpc-inference.ts <topic>');
          // Prefer server RDMA offers before assigning aliases. Transport is negotiated at load time.
          const selected = (await discoverRpcServers({ topic, timeoutMs: 5000 }))
              .sort((a, b) => Number(b.rdmaAvailable) - Number(a.rdmaAvailable) || a.url.localeCompare(b.url))
              .slice(0, 2);
          if (selected.length !== 2)
              throw new Error('Two idle RPC servers are required');
          const devices = getRpcDeviceMap(selected);
          if (!devices.length)
              throw new Error('The selected servers expose no devices');
          console.error(`▸ Using ${devices.length} devices across ${selected.length} servers`);
          const modelId = await loadModel({
              modelSrc: LLAMA_3_2_1B_INST_Q4_0,
              modelConfig: {
                  device: 'gpu',
                  'rpc-servers': selected.map((server) => server.url).join(','),
                  devices: devices.map((device) => device.alias).join(','),
                  'split-mode': 'layer',
                  'tensor-split': devices.map(() => '1').join(',')
              },
              onProgress: (progress) => {
                  process.stderr.write(`\r▸ Downloading ${progress.percentage.toFixed(0)}%`);
              }
          });
          try {
              process.stderr.write('\n');
              const run = completion({
                  modelId,
                  history: [{ role: 'user', content: 'Explain distributed inference in one sentence.' }],
                  stream: true
              });
              for await (const token of run.tokenStream)
                  process.stdout.write(token);
              process.stdout.write('\n');
          }
          finally {
              await unloadModel({ modelId });
          }
      }
      catch (error) {
          console.error('✖', error);
          process.exitCode = 1;
      }
      finally {
          await close();
      }
      ```
    </WrapCode>
  </Tab>

  <Tab value="ts" label="TypeScript">
    <WrapCode>
      ```ts file=<rootDir>/packages/sdk/examples/rpc-inference.ts title="rpc-inference.ts" lineNumbers
      // npx tsx rpc-inference.ts my-private-rpc-group
      import {
        discoverRpcServers,
        getRpcDeviceMap,
        loadModel,
        completion,
        unloadModel,
        close,
        LLAMA_3_2_1B_INST_Q4_0
      } from '@qvac/sdk'

      try {
        const topic = process.argv[2]
        if (!topic) throw new Error('Usage: rpc-inference.ts <topic>')
        // Prefer server RDMA offers before assigning aliases. Transport is negotiated at load time.
        const selected = (await discoverRpcServers({ topic, timeoutMs: 5000 }))
          .sort((a, b) => Number(b.rdmaAvailable) - Number(a.rdmaAvailable) || a.url.localeCompare(b.url))
          .slice(0, 2)
        if (selected.length !== 2) throw new Error('Two idle RPC servers are required')
        const devices = getRpcDeviceMap(selected)
        if (!devices.length) throw new Error('The selected servers expose no devices')
        console.error(`▸ Using ${devices.length} devices across ${selected.length} servers`)
        const modelId = await loadModel({
          modelSrc: LLAMA_3_2_1B_INST_Q4_0,
          modelConfig: {
            device: 'gpu',
            'rpc-servers': selected.map((server) => server.url).join(','),
            devices: devices.map((device) => device.alias).join(','),
            'split-mode': 'layer',
            'tensor-split': devices.map(() => '1').join(',')
          },
          onProgress: (progress) => {
            process.stderr.write(`\r▸ Downloading ${progress.percentage.toFixed(0)}%`)
          }
        })
        try {
          process.stderr.write('\n')
          const run = completion({
            modelId,
            history: [{ role: 'user', content: 'Explain distributed inference in one sentence.' }],
            stream: true
          })
          for await (const token of run.tokenStream) process.stdout.write(token)
          process.stdout.write('\n')
        } finally {
          await unloadModel({ modelId })
        }
      } catch (error) {
        console.error('✖', error)
        process.exitCode = 1
      } finally {
        await close()
      }
      ```
    </WrapCode>
  </Tab>
</Tabs>

<Callout type="success">
  All examples are self-contained and runnable. See the [JS/TS quickstart](/sdk/js-ts-sdk#quickstart) or [Python quickstart](/sdk/python-sdk#quickstart) for instructions.
</Callout>
