New: TranslatePsy-AfriSLM translates directly between 19 African languages, offline.
QVAC Logo
SDKAssess model fit
v0.20, not the current release

Assess model fit

Check whether a model is likely to fit in memory before downloading it.

Overview

assessModelFit() answers one question before anything is downloaded: is this model likely to fit in memory? It reads the model constant metadata together with a fresh memory sample, and returns a verdict for each candidate plus one for the set. For a single candidate it also runs the engine's own fitter, against the artifact where it is already on disk, and otherwise against the registry's weightless description of it — tens of KB, the tensor list with no data section.

It is advisory. It never downloads weights or loads a model, and it does not block loadModel(), reserve memory, pick a model for you, or make any claim about speed.

A catalog model constant or a path already on disk gets a verdict.

Functions

  1. assessModelFit() — describe the loads you are considering, as loadModel() takes them.
  2. downloadAsset() or loadModel() — fetch whichever candidate you picked.

For how to use each function, see SDK — API reference.

Parameters

assessModelFit({
  models: [
    {
      modelSrc: QWEN3_8B_INST_Q4_K_M,
      modelType: "llm",
      modelConfig: { ctx_size: 8192 },
    },
  ],
  execution: "sequential",
});

models

The candidates, at least one. They are assessed together, against a single budget. Each takes the same parameters as loadModel():

  • modelSrc — the model, as loadModel() takes it. A catalog constant resolves to the registry's weightless description of the artifact; a path already on disk is read directly.
  • modelType — optional. The engine that would run the load, inferred from modelSrc when omitted, as loadModel() infers it. Required only where the source names no engine.
  • modelConfig — optional. The config loadModel() would be given, including the companion model sources a compound load needs.

The engine's own plugin resolves that config, so the settings the fitter reads are the ones the load would run with, and the same object can be handed to loadModel() afterwards. The estimator reads its workload from the same config: ctx_size for a token context, duration_ms for an audio window.

execution

Either sequential or concurrent:

  • sequential: counts every model as resident but adds only the largest single working peak.
  • concurrent: adds every peak.

The mode describes what you intend to do so the numbers match your plan — the SDK does not schedule, serialize, or reserve anything on the strength of it.

Defaults to sequential.

policy

How much memory to withhold from the budget as headroom, left for the rest of the system:

  • interactive-v1: withholds 20% of the memory free at the time of the call, capped at 2 GiB on desktop and 1 GiB on mobile. Right now, it is the only accepted value and it applies by default, so passing it is optional.

Returns

The result answers for the whole set, and result.models repeats the same shape for each candidate. These are the fields to branch on.

verdict

VerdictMeaning
likely-fitsThe conservative upper bound is within the memory budget.
likely-too-largeEven the optimistic lower bound exceeds the budget.
unknownThe evidence does not support either claim.

evidence

What backs the verdict.

evidenceWhat it isCan say
native-fitThe engine's own fitter, run against the artifact on disk, or against the registry's weightless description of it.any verdict
calibrationA two-sided estimate from coefficients measured on this platform.any verdict
computed-onlyA floor from catalog facts alone — artifact bytes, plus the KV cache for llama.cpp models — reported as floorBytes.likely-too-large or unknown

native-fit is the strongest: it is the answer the loader itself would give on this machine, so it outranks calibration where both exist. It carries no estimate, because the fitter returns a plan rather than a byte range, so branch on the verdict rather than on bounds. It is reachable only for a single candidate, and only where every source the load names is a path on disk or has a registry description; a set of candidates, and a source that is not on disk and whose description cannot be fetched, keep the calibrated estimate.

Read evidence alongside the verdict, because the same unknown can mean different things:

  • Under calibration, the numbers came out too close to call.
  • Under computed-only, this platform has no measurements at all, so the answer says nothing about whether the model fits.

That difference matters once you act on the result. Hiding the likely-too-large candidates works the same way under all three, since that verdict is trustworthy either way. Telling a user that a model will fit needs native-fit or calibration; the computed floor never produces likely-fits.

basis

The memory the verdict was weighed against:

  • system-memory — device RAM and system-wide use. Desktop, and Android, whose low-memory killer acts system-wide.
  • process-memory — the app's own ceiling. iOS, where jetsam terminates an app on its own footprint against a limit well below device RAM.
  • device-memory — a discrete GPU's own memory, when the model would execute there.
  • device-budget — on Windows, the GPU memory the OS grants this process, since a card's readings there are per-process rather than device-wide.

Both device bases additionally require the system-memory budget to hold, because a GPU load is paid for in system RAM too.

device

Per model: where that load resolved to execute, gpu or cpu, after the device defaults this host applies for itself. Absent for an engine that expresses no placement.

Read it alongside evidence. The llama fitters read device memory alone and decline a cpu load, so one there carries no native-fit — its verdict comes from calibration where this platform has coefficients, and from the computed floor where it does not. The speech and voice fitters answer for a cpu load like any other.

reasons

Every unknown names its cause here — on the model when the cause is that model, on the result when it is the machine. A verdict that fell back because the fitter produced nothing says so, and names what the fitter reported. The ones you are most likely to meet:

  • No validated calibration for this platform, or for the GPU placement the model would use.
  • The model is not a catalog constant, so there is no resource profile to read.
  • The model's engine has no estimator, so nothing sizes the load beyond its weights.

Where verdicts are available

calibration evidence — and with it any likely-fits — exists only where coefficients have been measured on real hardware and validated against a held-out model. Today that is LLM workloads (llamacpp-completion, llamacpp-embedding) on desktop, except win32-arm64. Where a GPU would run the model, the platform also needs a fixture measured on that placement.

Everywhere else the assessment falls back to the computed floor: likely-too-large when the weights alone exceed the budget, unknown otherwise. That covers Android, iOS, win32-arm64, audio workloads, and every engine with no estimator yet.

native-fit is independent of calibration. Wherever the engine's own fitter runs — every platform that ships that engine's addon, mobile included — a single candidate on disk or with a published description gets the fitter's verdict, including likely-fits, before any coefficients are consulted.

The per-platform calibration matrix, which changes with every release, lives in packages/sdk/docs/assess-model-fit.md.

Load-time probe

Once the weights are on disk, loadModel() runs the same engine fitter against the real file and the resolved load settings before it loads. It is the stronger evidence once you have the file. Desktop runs it in a disposable child process. Mobile, which cannot spawn one, runs it on a worker thread.

The probe is advisory too. No verdict blocks or changes the load: does-not-fit is logged and the load runs unchanged, and a crash, timeout, or unsupported configuration resolves to no verdict. Verdicts appear on the SDK log stream as [advisory-fit:…] lines.

getLoadedModelInfo() returns the outcome as fitProbe:

  • verdict — fit, does-not-fit, or unknown. These are unhedged, unlike the verdicts above, because the probe measured this build against this file.
  • engine — the engine package whose fitter ran.
  • reason and message — why the probe reached that verdict.
  • plan — the placement it projected, where the fitter resolved one: nCtx, nGpuLayers, nGpuDevices.
  • projection — what the fitter measured, present on fit and does-not-fit alike. Engines do not all measure the same things, so every field is optional: deviceBytes and hostBytes for the peaks, weightsBytes, contextBytes and computeBytes within deviceBytes where the engine separates them, deviceFreeBytes and deviceTotalBytes, and report, the engine's own memory table.

The probe is on by default. Set the QVAC_ADVISORY_MODEL_FIT=0 environment variable to skip it, and fitProbe is then absent.

Examples

Assess before download

The following script asks which of five Qwen3 models this machine is likely to run at 8192 tokens of context, printing the budget basis, each verdict against the model's size on disk, and the reasons behind every unknown:

assess-model-fit.js
/**
 * Which models is this machine likely to run — asked before downloading any
 * weights. Nothing is fetched, loaded or reserved.
 *
 * The header prints the evidence the answer rests on: `system-memory` on a
 * CPU-only or integrated-GPU host, `device-memory` on a discrete card,
 * `device-budget` for the per-process allowance Windows grants on one. Which
 * of those applies depends on where the engine would put the model.
 *
 * `unknown` is a real answer, not an error: the evidence does not support a
 * call either way. Show it as "can't say", never as "no".
 */
import { assessModelFit, QWEN3_600M_INST_Q4, QWEN3_1_7B_INST_Q4, QWEN3_4B_INST_Q4_K_M, QWEN3_8B_INST_Q4_K_M, QWEN3_8_27B_MULTIMODAL_UD_Q8_K_XL } from '@qvac/sdk';
// Sizes the assessment the way it sizes the load: the estimator reads the
// context out of the config `loadModel` would be given.
const CONTEXT_TOKENS = 8192;
// A ladder ending well past what a laptop has, so one screen shows every verdict.
const CANDIDATES = [
    QWEN3_600M_INST_Q4,
    QWEN3_1_7B_INST_Q4,
    QWEN3_4B_INST_Q4_K_M,
    QWEN3_8B_INST_Q4_K_M,
    QWEN3_8_27B_MULTIMODAL_UD_Q8_K_XL
];
const VERDICT_MARK = {
    'likely-fits': '✔',
    'likely-too-large': '✖',
    unknown: '?'
};
function gib(bytes) {
    return `${(bytes / 1024 ** 3).toFixed(2)} GiB`;
}
try {
    const result = await assessModelFit({
        models: CANDIDATES.map((model) => ({
            modelSrc: model,
            modelType: 'llamacpp-completion',
            modelConfig: { ctx_size: CONTEXT_TOKENS }
        })),
        // Declared for aggregation only, not a scheduling instruction: 'sequential'
        // counts the largest operation peak, 'concurrent' counts one per model.
        execution: 'sequential',
        policy: 'interactive-v1'
    });
    console.log(`▸ Budget basis: ${result.basis}`);
    if (result.budget) {
        console.log(`    ${gib(result.budget.availableAfterReserveBytes)} budget` +
            ` (${gib(result.budget.totalBytes)} total,` +
            ` ${gib(result.budget.usedBytes)} in use,` +
            ` ${gib(result.budget.availableBytes)} free,` +
            ` ${gib(result.budget.reservedBytes)} held back)`);
    }
    // Says which device the verdicts assume, when the model would run on a GPU.
    const placement = result.assumptions.find((line) => line.includes('assumed to execute on'));
    if (placement)
        console.log(`▸ ${placement}`);
    console.log(`\n▸ At ${CONTEXT_TOKENS} tokens of context`);
    for (const [index, model] of result.models.entries()) {
        const mark = VERDICT_MARK[model.verdict];
        const size = gib(CANDIDATES[index].expectedSize).padStart(9);
        // A calibrated row has a two-sided estimate; a computed-only row has the floor.
        const needs = model.estimate
            ? `needs ${gib(model.estimate.upperBoundBytes)}`
            : model.evidence === 'computed-only' && model.floorBytes
                ? `at least ${gib(model.floorBytes)}`
                : '';
        console.log(`  ${mark} ${model.name.padEnd(46)} ${size} on disk  ${needs}`);
        // Present on every `unknown`, and worth surfacing: it names what is missing.
        if (model.verdict === 'unknown') {
            for (const reason of model.reasons)
                console.log(`      ${reason}`);
        }
    }
    console.log(`\n▸ All five together, sequentially: ${result.verdict}`);
    console.log('▸ Advisory only — this does not gate loadModel or reserve anything.');
    process.exit(0);
}
catch (error) {
    console.error('✖', error);
    process.exit(1);
}

Load-time probe

The following script loads Qwen3.5 0.8B at 4k context, which the fitter projects to fit, and reprints the [advisory-fit:…] verdicts from the SDK log stream. Set QVAC_FIT_DEMO_ATTEMPT_OVERSIZED=1 to also load gpt-oss-20B at 128k context with an f32 KV cache, which it projects not to fit and loads anyway:

advisory-model-fit.js
/**
 * Advisory model fit check.
 *
 * Before a load, the SDK runs the fitter belonging to the engine that would
 * run it and projects whether the exact configuration it is about to load will
 * fit in device memory. Desktop runs it in a disposable Bare child; mobile,
 * which cannot spawn one, runs it on a worker thread.
 *
 * The result is ADVISORY. It never blocks a load. `does-not-fit` is logged and
 * the ordinary load path runs unchanged. Crashes, timeouts, malformed
 * responses, unsupported configurations, and internal errors all resolve to
 * "no evidence" and are equally non-blocking. No verdict changes the load;
 * `getLoadedModelInfo` returns it as `fitProbe`.
 *
 * The verdict is emitted on the SDK server log stream, not to stdout, so this
 * example subscribes to `loggingStream({ id: SDK_LOG_ID })` and reprints the
 * `[advisory-fit:…]` lines.
 *
 * ---------------------------------------------------------------------------
 * What fits on this machine (Apple M4 Pro, 24 GiB unified memory)
 * ---------------------------------------------------------------------------
 *
 * Measured at the default 1024 MiB margin. The fitter budgets against what the
 * machine can actually keep resident
 * (total − wired − compressor: 17.4 GiB on this machine at idle), not the raw
 * RAM figure.
 *
 *   PROJECTED TO FIT — all layers on GPU
 *     Qwen3.5 0.8B  Q4_K_M   0.5 GiB  @   4k ctx
 *     gpt-oss-20B   Q4_K_M  10.8 GiB  @  32k ctx
 *     gte-large     fp16     0.6 GiB  (embedding, context pinned to 512)
 *
 *   PROJECTED NOT TO FIT — with `gpu_layers: 99` pinned
 *     gpt-oss-20B   Q4_K_M  10.8 GiB  @ 128k ctx with an f32 KV cache ← below
 *
 * These rows were measured with `gpu_layers` pinned at 99. Loads no longer pin
 * the layer count, and the `does-not-fit` row has not been re-measured
 * unpinned.
 *
 * gpt-oss-20B at 128k context FITS with the default KV cache and DOES NOT FIT
 * once `cache-type-k`/`cache-type-v` are set to `f32`. Same model, same
 * context, same machine — the verdict tracks the configuration, not the file
 * size.
 *
 * Two boundaries worth understanding when reading verdicts:
 *
 * 1. The verdict answers for a PLACEMENT, not a model. With `gpu_layers`
 *    unset, the fitter may move layers to the CPU side; setting it pins the
 *    layer count, and `does-not-fit` then means "not at this placement".
 * 2. A `does-not-fit` configuration can still RUN on macOS when the OS
 *    compresses and pages hard enough — and a `fits` configuration right at
 *    the boundary can still fail at first decode under memory pressure.
 *    Prediction cannot separate those cases from a snapshot; the addon-side
 *    probe decode (QVAC-24114) is the runtime check that catches the
 *    remainder. The verdict here is the honest working-set budget, and the
 *    measured failure modes punish over-committing, so treat `does-not-fit`
 *    as "expect degradation or decode failure", not "the load will error".
 */
import { completion, loadModel, unloadModel, loggingStream, SDK_LOG_ID, QWEN3_5_0_8B_MULTIMODAL_Q4_K_M, GPT_OSS_20B_INST_Q4_K_M } from '@qvac/sdk';
// The oversized load below is expected to be reported as `does-not-fit` and
// then attempted anyway, because the check is advisory. It really does try to
// allocate ~11 GiB, so it stays opt-in.
const ATTEMPT_OVERSIZED = process.env['QVAC_FIT_DEMO_ATTEMPT_OVERSIZED'] === '1';
// Reprint the worker's advisory verdicts. They arrive on the SDK server log
// stream; everything else on that stream is filtered out to keep this readable.
//
// Called once per phase rather than once for the process: a subscription
// currently stops delivering after any `unloadModel`, so a single one would go
// silent before the second verdict. Resubscribing after the unload works.
function watchVerdicts() {
    void (async () => {
        for await (const log of loggingStream({ id: SDK_LOG_ID })) {
            if (log.message.includes('[advisory-fit:')) {
                console.log(`▸ [${log.level.toUpperCase()}] ${log.message}`);
            }
        }
    })().catch(() => {
        // Stream terminated — normal on shutdown.
    });
}
watchVerdicts();
try {
    // 1. A load the fitter projects to fit. The verdict carries the plan it
    //    projected: resolved context, offloaded layers, and GPU device count.
    console.log('▸ Loading Qwen3.5 0.8B @ 4k — expected verdict: projected to fit');
    const smallModelId = await loadModel({
        modelSrc: QWEN3_5_0_8B_MULTIMODAL_Q4_K_M,
        modelConfig: { ctx_size: 4096 }
    });
    console.log(`▸ Loaded ${smallModelId}\n`);
    const result = completion({
        modelId: smallModelId,
        history: [{ role: 'user', content: 'Say hello in five words.' }],
        stream: false,
        generationParams: { predict: 48 }
    });
    const final = await result.final;
    console.log(`▸ Completion still works normally: ${final.contentText.trim().slice(0, 120)}\n`);
    // Unloaded before the next phase, so the second verdict is measured on an
    // idle machine and stays comparable to the fixture tables above. Leaving it
    // loaded would shift the verdict: the check reserves resident weight bytes
    // through the fit margin, and the fitter also sees system-wide wired memory.
    await unloadModel({ modelId: smallModelId, clearStorage: false });
    watchVerdicts();
    // 2. A load the fitter projects NOT to fit. The point of this example is that
    //    the SDK reports the verdict and then loads anyway — the check is
    //    evidence, not admission control.
    if (!ATTEMPT_OVERSIZED) {
        console.log('▸ Skipping the oversized gpt-oss-20B load.');
        console.log('▸ Set QVAC_FIT_DEMO_ATTEMPT_OVERSIZED=1 to let it run and watch the');
        console.log('  load proceed past a `does-not-fit` verdict (allocates ~11 GiB).');
    }
    else {
        console.log('▸ Loading gpt-oss-20B @ 128k with an f32 KV cache');
        console.log('▸ Expected verdict: projected NOT to fit');
        console.log('▸ The load is attempted regardless. That is the fail-open contract:');
        console.log('  the verdict is evidence, not admission control.\n');
        const bigModelId = await loadModel({
            modelSrc: GPT_OSS_20B_INST_Q4_K_M,
            modelConfig: {
                ctx_size: 131072,
                'cache-type-k': 'f32',
                'cache-type-v': 'f32'
            }
        });
        console.log(`▸ Load returned ${bigModelId} — the advisory verdict did not block it`);
        // Loading is not the same as being usable: a model can load and then fail
        // at decode time. Gemma 4 31B does exactly that on this machine. So run a
        // real completion and report throughput rather than trusting the load.
        try {
            const check = completion({
                modelId: bigModelId,
                history: [{ role: 'user', content: 'Name three colours. Answer briefly.' }],
                stream: false,
                generationParams: { predict: 40 }
            });
            const checkFinal = await check.final;
            console.log(`▸ ...and it actually runs: ${checkFinal.stats?.tokensPerSecond?.toFixed(1) ?? '?'} tok/s ` +
                `— the OS compressed and paged its way past the working-set budget`);
        }
        catch (inferenceError) {
            console.log(`▸ ...but it cannot run: ${inferenceError instanceof Error ? inferenceError.message : String(inferenceError)}`);
            console.log('▸ The verdict was right, and `loadModel` succeeding did not mean usable.');
        }
        await unloadModel({ modelId: bigModelId, clearStorage: false });
    }
}
catch (error) {
    // A failing load here is the native loader's own error, not the fit check.
    // The check never throws and never converts a verdict into a load failure.
    console.error('✖', error instanceof Error ? error.message : error);
    process.exitCode = 1;
}
// The log subscription is an open stream and would otherwise keep the process
// alive after the work is done.
process.exit(process.exitCode ?? 0);

The Python client supports this capability through the same worker. A dedicated Python example is not yet published — see the Python SDK for the API surface.

Tip: all examples throughout this documentation are self-contained and runnable. For instructions on how to run them, see the JS/TS quickstart or the Python quickstart.

On this page

Ask anything about QVAC.