---
title: "Profiler"
canonical: https://docs.qvac.tether.io/sdk/v0.19/runtime/profiler/
collection: "SDK"
package: "@qvac/sdk"
line: v0.19
current_line: false
---

# Profiler (/sdk/v0.19/runtime/profiler)



## Overview

`@qvac/sdk` npm package exposes a `profiler` object that you can import and use to measure and analyze how long SDK operations take in your application. You can enable profiling in two ways:

* [**Global:**](#enable-global) call `profiler.enable()` to profile all subsequent operations until `profiler.disable()` is called.
* [**Per-call:**](#enable-per-call) pass `{ profiling: { enabled: true } }` in the options of an individual function call.

You can also enable profiling globally and use a per-call override to [opt out](#opt-out) of specific calls by passing `{ profiling: { enabled: false } }`. Data collected through both modes is stored in and exported from the same `profiler` singleton. See [Support](#support) for the list of SDK operations that support profiling.

<Callout type="info">
  `profiler` is a process-wide singleton — its state persists until the process exits.
</Callout>

<Callout type="info">
  **Python:** the process-wide `profiler` object is JS/TS-only and intentionally not ported. In the Python SDK, profile a single call with `profiled_call()` and the `__profiling` envelope helpers from `tetherto.qvac_sdk.profiling`, then aggregate the returned `ProfilingReport`s yourself.
</Callout>

## `profiler`

The `profiler` object provides the following methods:

1. `profiler.enable(options?)` — enable profiling globally
2. `profiler.isEnabled()` — return whether profiling is enabled
3. `profiler.onRecord(callback)` — subscribe to profiling events in real time
4. Export collected data using any of the following:
   * `profiler.exportSummary()` — return a high-level summary string
   * `profiler.exportTable()` — return a detailed table of aggregated metrics
   * `profiler.exportJSON(options?)` — return a full export as structured JSON
5. `profiler.disable()` — disable profiling
6. `profiler.clear()` — clear collected data

See [API — `profiler`](/sdk/v0.19/reference/api#profiler) for the complete reference, including the full metrics catalog for each export function and detailed usage.

## Enable profiling

By default, profiling is disabled. You can enable it [globally](#enable-global) or [per-call](#enable-per-call). If you enable profiling globally, you can [opt out](#opt-out) of individual function calls.

### Global

Enable profiling globally by calling `profiler.enable()`:

```ts
import { profiler } from "@qvac/sdk";

profiler.enable({
  mode: "verbose",                 // "summary" (default) | "verbose"
  includeServerBreakdown: true,    // include server-side timing in responses
  includeResourceGauges: true,     // attach one local resource sample per operation
  operationFilters: ["completion"], // only profile these operations (empty = all)
});
```

### Per-call

To profile a single call, pass the `profiling` option when invoking the function:

```ts
await embed(
  { modelId, text: "hello" },
  {
    profiling: {
      enabled: true,
      includeServerBreakdown: true,
      includeResourceGauges: true,
      mode: "verbose",
    },
  },
);
```

See [Support](#support) for the list of functions that accept the `profiling` option.

### Resource gauges

Resource gauges are opt-in. Set `includeResourceGauges: true` globally or for
one call to attach a timestamped local CPU, memory, and GPU sample to each
profiled operation event:

```ts
profiler.enable({
  mode: "verbose",
  includeResourceGauges: true,
});

// Run SDK operations, then inspect recentEvents[].resources.
const profile = profiler.exportJSON();
```

The gauges reuse the `ResourceMetric<T>` status, source, and scope semantics
from `getSystemResources`. Unsupported metrics stay explicit rather than
becoming zero. `resources.sampledAt` uses the same monotonic clock as the
profiling event's `ts`, so the timestamps can be compared directly.
Samples are emitted through `profiler.onRecord`; they appear
in `exportJSON().recentEvents` only in `verbose` mode. Enabling gauges in
`summary` mode still pays the sampling cost without retaining the samples.
Profiling-disabled operations and calls that do not opt in perform no resource
sampling. Resource gauges are operation-scoped diagnostics returned to the
requesting SDK, not exported telemetry, and do not affect model admission.
Enabling them adds one CPU query and one query per GPU to each profiled
operation's response path. If the worker resource collector is not initialized,
the event omits the resource block.

### Opt-out

If you enabled profiling [globally](#enable-global), you can pass the `profiling` option to opt out of a specific call:

```ts
await embed(
  { modelId, text: "hello" },
  { profiling: { enabled: false } },
);
```

## Support

The following SDK operations support profiling:

[`completion()`](/sdk/v0.19/reference/api#completion) | [`downloadAsset()`](/sdk/v0.19/reference/api#downloadasset) | [`embed()`](/sdk/v0.19/reference/api#embed) | [`invokePlugin()`](/sdk/v0.19/reference/api#invokeplugin) | [`invokePluginStream()`](/sdk/v0.19/reference/api#invokepluginstream) | [`loadModel()`](/sdk/v0.19/reference/api#loadmodel) | [`ocr()`](/sdk/v0.19/reference/api#ocr) | [`ragChunk()`](/sdk/v0.19/reference/api#ragchunk) | [`ragCloseWorkspace()`](/sdk/v0.19/reference/api#ragcloseworkspace) | [`ragDeleteEmbeddings()`](/sdk/v0.19/reference/api#ragdeleteembeddings) | [`ragDeleteWorkspace()`](/sdk/v0.19/reference/api#ragdeleteworkspace) | [`ragIngest()`](/sdk/v0.19/reference/api#ragingest) | [`ragListWorkspaces()`](/sdk/v0.19/reference/api#raglistworkspaces) | [`ragReindex()`](/sdk/v0.19/reference/api#ragreindex) | [`ragSaveEmbeddings()`](/sdk/v0.19/reference/api#ragsaveembeddings) | [`ragSearch()`](/sdk/v0.19/reference/api#ragsearch) | [`textToSpeech()`](/sdk/v0.19/reference/api#texttospeech) | [`transcribe()`](/sdk/v0.19/reference/api#transcribe) | [`transcribeStream()`](/sdk/v0.19/reference/api#transcribestream) | [`translate()`](/sdk/v0.19/reference/api#translate)

## Examples

### Global

The following script enables profiling globally, loads a model, runs a completion, and exports timing data while also capturing resource gauges and backend diagnostics:

<Tabs>
  <Tab value="js" label="JavaScript" default>
    <WrapCode>
      ```js file=<rootDir>/packages/sdk/dist/examples/profiling/basic.js title="profiling-basic.js" lineNumbers
      import { completion, loadModel, unloadModel, LLAMA_3_2_1B_INST_Q4_0, profiler } from '@qvac/sdk';
      try {
          // Enable profiling globally
          profiler.enable({
              mode: 'verbose',
              includeServerBreakdown: true,
              includeResourceGauges: true
          });
          console.log('▸ Profiler enabled:', profiler.isEnabled());
          const modelId = await loadModel({
              modelSrc: LLAMA_3_2_1B_INST_Q4_0,
              onProgress: (p) => {
                  const mb = (n) => (n / 1e6).toFixed(1);
                  const line = `▸ Downloading ${p.percentage.toFixed(0)}% (${mb(p.downloaded)}/${mb(p.total)} MB)`;
                  process.stderr.write(process.stderr.isTTY ? `\r${line}` : `${line}\n`);
                  if (p.percentage >= 100)
                      process.stderr.write('\n');
              }
          });
          console.log('▸ Model loaded:', modelId);
          console.log('\n▸ Running completion...');
          const result = completion({
              modelId,
              history: [{ role: 'user', content: 'Say hello in one sentence.' }],
              stream: true
          });
          for await (const token of result.tokenStream) {
              process.stdout.write(token);
          }
          console.log();
          await unloadModel({ modelId });
          // Export profiling data
          console.log('\n▸ Profiler Summary');
          console.log(profiler.exportSummary());
          console.log('\n▸ Profiler Table');
          console.log(profiler.exportTable());
          const json = profiler.exportJSON();
          console.log('\n▸ Load Model Metrics');
          // Filter for operation-level event (kind: "handler"), not RPC phase events
          const loadModelEvent = json.recentEvents?.find((e) => e.op === 'loadModel' && e.kind === 'handler');
          if (loadModelEvent) {
              const tags = loadModelEvent.tags ?? {};
              const gauges = loadModelEvent.gauges ?? {};
              console.log('  sourceType:', tags['sourceType'] ?? '(not set)');
              console.log('  cacheHit:', tags['cacheHit'] ?? '(not set)');
              console.log('  totalLoadTime:', gauges['totalLoadTime'], 'ms');
              console.log('  modelInitializationTime:', gauges['modelInitializationTime'], 'ms');
              if (tags['cacheHit'] !== 'true') {
                  console.log('  downloadTime:', gauges['downloadTime'] ?? '(cached)', 'ms');
                  console.log('  totalBytesDownloaded:', gauges['totalBytesDownloaded'] ?? '(cached)');
                  console.log('  downloadSpeedBps:', gauges['downloadSpeedBps'] ?? '(cached)');
              }
              else {
                  console.log('  (download metrics omitted - cache hit)');
              }
              if (gauges['checksumValidationTime'] !== undefined) {
                  console.log('  checksumValidationTime:', gauges['checksumValidationTime'], 'ms');
              }
          }
          else {
              console.log('  (no loadModel handler event captured)');
              // Debug: show what ops are available
              const ops = [...new Set(json.recentEvents?.map((e) => `${e.op}:${e.kind}`) ?? [])];
              console.log('  Available ops:', ops.join(', '));
          }
          const resourceEvent = json.recentEvents?.filter((event) => event.resources).at(-1);
          console.log('\n▸ Resource Gauges');
          if (resourceEvent?.resources) {
              const resources = resourceEvent.resources;
              console.log('  op:', resourceEvent.op);
              console.log('  cpu:', resources.cpu.status === 'supported'
                  ? `${(resources.cpu.value * 100).toFixed(1)}%`
                  : resources.cpu.status);
              console.log('  memoryUsed:', resources.memory.usedBytes.status === 'supported'
                  ? `${(resources.memory.usedBytes.value / 1024 ** 3).toFixed(2)} GiB`
                  : resources.memory.usedBytes.status);
              console.log('  gpuCount:', resources.gpus.status === 'supported' ? resources.gpus.value.length : resources.gpus.status);
          }
          else {
              console.log('  (no resource gauges reported)');
          }
          const backendEvent = json.recentEvents?.filter((event) => event.backend).at(-1);
          console.log('\n▸ Backend Diagnostics');
          if (backendEvent?.backend) {
              const backend = backendEvent.backend;
              console.log('  op:', backendEvent.op);
              console.log('  selectedBackend:', backend.selectedBackend);
              console.log('  selectedDevice:', backend.selectedDevice);
              console.log('  graphicsApi:', backend.graphicsApi ?? '(not reported)');
              console.log('  fallback:', backend.fallback?.reason ?? '(none)');
          }
          else {
              console.log('  (no backend diagnostics reported for this operation)');
          }
          console.log('\n▸ Profiler JSON (structure)');
          console.log('  aggregates:', Object.keys(json.aggregates).length, 'metrics');
          console.log('  recentEvents:', json.recentEvents?.length ?? 0, 'events');
          console.log('  config:', json.config);
          // Disable profiling
          profiler.disable();
          console.log('\n▸ Profiler disabled:', !profiler.isEnabled());
      }
      catch (error) {
          console.error('✖', error);
          process.exit(1);
      }
      ```
    </WrapCode>
  </Tab>

  <Tab value="ts" label="TypeScript">
    <WrapCode>
      ```ts file=<rootDir>/packages/sdk/examples/profiling/basic.ts title="profiling-basic.ts" lineNumbers
      import { completion, loadModel, unloadModel, LLAMA_3_2_1B_INST_Q4_0, profiler } from '@qvac/sdk'

      try {
        // Enable profiling globally
        profiler.enable({
          mode: 'verbose',
          includeServerBreakdown: true,
          includeResourceGauges: true
        })
        console.log('▸ Profiler enabled:', profiler.isEnabled())

        const modelId = await loadModel({
          modelSrc: LLAMA_3_2_1B_INST_Q4_0,
          onProgress: (p) => {
            const mb = (n: number) => (n / 1e6).toFixed(1)
            const line = `▸ Downloading ${p.percentage.toFixed(0)}% (${mb(p.downloaded)}/${mb(p.total)} MB)`
            process.stderr.write(process.stderr.isTTY ? `\r${line}` : `${line}\n`)
            if (p.percentage >= 100) process.stderr.write('\n')
          }
        })
        console.log('▸ Model loaded:', modelId)

        console.log('\n▸ Running completion...')
        const result = completion({
          modelId,
          history: [{ role: 'user', content: 'Say hello in one sentence.' }],
          stream: true
        })

        for await (const token of result.tokenStream) {
          process.stdout.write(token)
        }
        console.log()

        await unloadModel({ modelId })

        // Export profiling data
        console.log('\n▸ Profiler Summary')
        console.log(profiler.exportSummary())

        console.log('\n▸ Profiler Table')
        console.log(profiler.exportTable())

        const json = profiler.exportJSON()
        console.log('\n▸ Load Model Metrics')
        // Filter for operation-level event (kind: "handler"), not RPC phase events
        const loadModelEvent = json.recentEvents?.find(
          (e) => e.op === 'loadModel' && e.kind === 'handler'
        )
        if (loadModelEvent) {
          const tags = loadModelEvent.tags ?? {}
          const gauges = loadModelEvent.gauges ?? {}
          console.log('  sourceType:', tags['sourceType'] ?? '(not set)')
          console.log('  cacheHit:', tags['cacheHit'] ?? '(not set)')
          console.log('  totalLoadTime:', gauges['totalLoadTime'], 'ms')
          console.log('  modelInitializationTime:', gauges['modelInitializationTime'], 'ms')
          if (tags['cacheHit'] !== 'true') {
            console.log('  downloadTime:', gauges['downloadTime'] ?? '(cached)', 'ms')
            console.log('  totalBytesDownloaded:', gauges['totalBytesDownloaded'] ?? '(cached)')
            console.log('  downloadSpeedBps:', gauges['downloadSpeedBps'] ?? '(cached)')
          } else {
            console.log('  (download metrics omitted - cache hit)')
          }
          if (gauges['checksumValidationTime'] !== undefined) {
            console.log('  checksumValidationTime:', gauges['checksumValidationTime'], 'ms')
          }
        } else {
          console.log('  (no loadModel handler event captured)')
          // Debug: show what ops are available
          const ops = [...new Set(json.recentEvents?.map((e) => `${e.op}:${e.kind}`) ?? [])]
          console.log('  Available ops:', ops.join(', '))
        }

        const resourceEvent = json.recentEvents?.filter((event) => event.resources).at(-1)
        console.log('\n▸ Resource Gauges')
        if (resourceEvent?.resources) {
          const resources = resourceEvent.resources
          console.log('  op:', resourceEvent.op)
          console.log(
            '  cpu:',
            resources.cpu.status === 'supported'
              ? `${(resources.cpu.value * 100).toFixed(1)}%`
              : resources.cpu.status
          )
          console.log(
            '  memoryUsed:',
            resources.memory.usedBytes.status === 'supported'
              ? `${(resources.memory.usedBytes.value / 1024 ** 3).toFixed(2)} GiB`
              : resources.memory.usedBytes.status
          )
          console.log(
            '  gpuCount:',
            resources.gpus.status === 'supported' ? resources.gpus.value.length : resources.gpus.status
          )
        } else {
          console.log('  (no resource gauges reported)')
        }

        const backendEvent = json.recentEvents?.filter((event) => event.backend).at(-1)
        console.log('\n▸ Backend Diagnostics')
        if (backendEvent?.backend) {
          const backend = backendEvent.backend
          console.log('  op:', backendEvent.op)
          console.log('  selectedBackend:', backend.selectedBackend)
          console.log('  selectedDevice:', backend.selectedDevice)
          console.log('  graphicsApi:', backend.graphicsApi ?? '(not reported)')
          console.log('  fallback:', backend.fallback?.reason ?? '(none)')
        } else {
          console.log('  (no backend diagnostics reported for this operation)')
        }

        console.log('\n▸ Profiler JSON (structure)')
        console.log('  aggregates:', Object.keys(json.aggregates).length, 'metrics')
        console.log('  recentEvents:', json.recentEvents?.length ?? 0, 'events')
        console.log('  config:', json.config)

        // Disable profiling
        profiler.disable()
        console.log('\n▸ Profiler disabled:', !profiler.isEnabled())
      } catch (error) {
        console.error('✖', error)
        process.exit(1)
      }
      ```
    </WrapCode>
  </Tab>
</Tabs>

### Per-call

The following script keeps the profiler disabled globally and selectively profiles individual `embed()` calls using the per-call option:

<Tabs>
  <Tab value="js" label="JavaScript" default>
    <WrapCode>
      ```js file=<rootDir>/packages/sdk/dist/examples/profiling/per-call.js title="profiling-per-call.js" lineNumbers
      import { embed, loadModel, unloadModel, GTE_LARGE_FP16, profiler } from '@qvac/sdk';
      try {
          profiler.disable();
          console.log('▸ Profiler globally enabled:', profiler.isEnabled());
          const modelId = await loadModel({
              modelSrc: GTE_LARGE_FP16,
              onProgress: (p) => {
                  const mb = (n) => (n / 1e6).toFixed(1);
                  const line = `▸ Downloading ${p.percentage.toFixed(0)}% (${mb(p.downloaded)}/${mb(p.total)} MB)`;
                  process.stderr.write(process.stderr.isTTY ? `\r${line}` : `${line}\n`);
                  if (p.percentage >= 100)
                      process.stderr.write('\n');
              }
          });
          console.log('▸ Model loaded:', modelId);
          console.log('\n▸ Embed with per-call profiling');
          const { embedding: embedding1 } = await embed({ modelId, text: 'Profile this specific call' }, { profiling: { enabled: true, includeServerBreakdown: true } });
          console.log('▸ Embedding dimensions:', embedding1.length);
          console.log('\n▸ Embed without profiling');
          const { embedding: embedding2 } = await embed({
              modelId,
              text: 'This call is not profiled'
          });
          console.log('▸ Embedding dimensions:', embedding2.length);
          console.log('\n▸ Embed with profiling explicitly disabled');
          const { embedding: embedding3 } = await embed({ modelId, text: 'Profiling explicitly disabled for this call' }, { profiling: { enabled: false } });
          console.log('▸ Embedding dimensions:', embedding3.length);
          await unloadModel({ modelId });
          console.log('\n▸ Profiler Summary (per-call data only)');
          console.log(profiler.exportSummary());
      }
      catch (error) {
          console.error('✖', error);
          process.exit(1);
      }
      ```
    </WrapCode>
  </Tab>

  <Tab value="ts" label="TypeScript">
    <WrapCode>
      ```ts file=<rootDir>/packages/sdk/examples/profiling/per-call.ts title="profiling-per-call.ts" lineNumbers
      import { embed, loadModel, unloadModel, GTE_LARGE_FP16, profiler } from '@qvac/sdk'

      try {
        profiler.disable()
        console.log('▸ Profiler globally enabled:', profiler.isEnabled())

        const modelId = await loadModel({
          modelSrc: GTE_LARGE_FP16,
          onProgress: (p) => {
            const mb = (n: number) => (n / 1e6).toFixed(1)
            const line = `▸ Downloading ${p.percentage.toFixed(0)}% (${mb(p.downloaded)}/${mb(p.total)} MB)`
            process.stderr.write(process.stderr.isTTY ? `\r${line}` : `${line}\n`)
            if (p.percentage >= 100) process.stderr.write('\n')
          }
        })
        console.log('▸ Model loaded:', modelId)

        console.log('\n▸ Embed with per-call profiling')
        const { embedding: embedding1 } = await embed(
          { modelId, text: 'Profile this specific call' },
          { profiling: { enabled: true, includeServerBreakdown: true } }
        )
        console.log('▸ Embedding dimensions:', embedding1.length)

        console.log('\n▸ Embed without profiling')
        const { embedding: embedding2 } = await embed({
          modelId,
          text: 'This call is not profiled'
        })
        console.log('▸ Embedding dimensions:', embedding2.length)

        console.log('\n▸ Embed with profiling explicitly disabled')
        const { embedding: embedding3 } = await embed(
          { modelId, text: 'Profiling explicitly disabled for this call' },
          { profiling: { enabled: false } }
        )
        console.log('▸ Embedding dimensions:', embedding3.length)

        await unloadModel({ modelId })

        console.log('\n▸ Profiler Summary (per-call data only)')
        console.log(profiler.exportSummary())
      } catch (error) {
        console.error('✖', error)
        process.exit(1)
      }
      ```
    </WrapCode>
  </Tab>
</Tabs>

<Callout type="success">
  **Tip:** all examples throughout this documentation are self-contained and runnable. For instructions on how to run them, see the [JS/TS quickstart](/sdk/v0.19/js-ts-sdk#quickstart) or the [Python quickstart](/sdk/v0.19/python-sdk#quickstart).
</Callout>
