---
title: "Download lifecycle"
canonical: https://docs.qvac.tether.io/sdk/models/download-lifecycle/
collection: "SDK"
package: "@qvac/sdk"
line: v0.21
current_line: true
---

# Download lifecycle (/sdk/models/download-lifecycle)



## Overview

Downloads in QVAC are *resumable by default*. When you download an asset via [`downloadAsset()`](/sdk/reference/api#downloadasset) or [`loadModel()`](/sdk/reference/api#loadmodel)), the SDK writes partial files to disk so the next run can continue from where it left off. The progress callback provides a `downloadKey` that identifies the underlying transfer (useful for dedup and cache identification), but cancellation is targeted by `requestId`.

Both `downloadAsset()` and `loadModel()` return a decorated promise (`Promise<string> & { requestId: string }`) that exposes a synchronous `requestId` field, so you can wire a stop button to a specific in-flight call without waiting for the first progress event. See [Cancel a specific call by `requestId`](#cancel-a-specific-call-by-requestid) below.

## Functions

1. [`downloadAsset()`](/sdk/reference/api#downloadasset) or [`loadModel()`](/sdk/reference/api#loadmodel) — with `onProgress` for progress tracking; both return a decorated promise that exposes `op.requestId` synchronously.
2. [`cancel()`](/sdk/reference/api#cancel) — either:
   * `cancel({ requestId: op.requestId })` — pause this specific call (preserves the partial file for automatic resume on the next run).
   * `cancel({ requestId: op.requestId, clearCache: true })` — discard the partial file along with the cancel.
   * `cancel({ modelId })` — broad sweep that cancels every in-flight request on the given model, including non-download ops. See [Cancellation — broad cancel by `modelId`](/sdk/runtime/cancellation#broad-cancel-by-modelid-escape-hatch).

For how to use each function, see [SDK — API reference](/sdk/reference/api/).

## Prepare catalog models for offline use

Use `downloadAsset()` to provision a catalog model without loading it into memory. A later `loadModel()` call with the same catalog constant checks the configured cache first, validates the cached files against bundled size and checksum metadata, and can load them without contacting the registry.

```ts
import { downloadAsset, loadModel, QWEN3_4B_INST_Q4_K_M } from "@qvac/sdk";

// Run this while the registry is reachable.
await downloadAsset({ assetSrc: QWEN3_4B_INST_Q4_K_M });

// This can run offline once the download has completed.
const modelId = await loadModel({ modelSrc: QWEN3_4B_INST_Q4_K_M });
```

The SDK instance that performs each call must use the same [`cacheDirectory`](/sdk/configuration#options). The initial download still requires registry access; this workflow prepares an application for later offline startup. To load a catalog model from an alternate source when the registry itself is unreachable, see [Fall back to an alternate source](#fall-back-to-an-alternate-source).

If an application manages its own model files, pass a local path and an explicit model type instead:

```ts
const modelId = await loadModel({
  modelSrc: "/opt/models/model.gguf",
  modelType: "llamacpp-completion",
});
```

Local-path loading does not associate the file with a catalog constant or validate it against catalog checksum metadata. Applications using this path are responsible for provisioning and integrity checks.

## Fall back to an alternate source

On some networks the registry is not reliably reachable, and a `loadModel()` call using a catalog constant can fail before the model arrives. Pass `fallbackSrc` — an HTTP URL or a local file path — to load the same model from an alternate source when the registry download does not succeed:

```ts
import { loadModel, QWEN3_4B_INST_Q4_K_M } from "@qvac/sdk";

const modelId = await loadModel({
  modelSrc: QWEN3_4B_INST_Q4_K_M,
  fallbackSrc: "https://mirror.example.com/qwen3-4b-instruct-q4_k_m.gguf",
});
```

The model loaded from `fallbackSrc` is validated against the catalog model's checksum before use, so an alternate source is trusted to the same degree as the registry copy. Download progress for the fallback is reported through the same `onProgress` callback.

`fallbackSrc` is supported only when `modelSrc` is a built-in catalog constant — the constant supplies the checksum to validate against. A local-path or URL `modelSrc` has no catalog checksum, so `fallbackSrc` is rejected for those. When a catalog model cannot be downloaded and no `fallbackSrc` was given, the error suggests supplying one.

## Source trust and transport

Model sources fall into three trust tiers:

* **Catalog / `registry://` (and `fallbackSrc`)** — verified against the checksum bundled with the catalog constant. A `fallbackSrc` is validated against that same checksum, so it is trusted to the same degree as the registry copy.
* **Hugging Face HTTP URLs** (`huggingface.co` / `hf.co`) — **Hub-attested**. The SDK reads the file's SHA-256 from the Hub (its `X-Linked-Etag`) and verifies the downloaded bytes against it. This covers single-file, sharded, and archive (`.tar`/`.tar.gz`) downloads. If a Hugging Face URL exposes no SHA-256 (for example a small non-LFS file), the download proceeds unverified with a warning — unless `requireHttpChecksum` is enabled (see below).
* **Other HTTP(S) URLs (bring-your-own)** — downloaded **as-is and unverified**. There is no trusted checksum to check them against. Plaintext `http://` on any host is accepted and an `https://` URL that redirects to `http://` is followed — pointing at your own model server (on any domain) works unchanged. Every download whose integrity cannot be verified (bring-your-own HTTP, or a Hugging Face file with no published SHA-256) logs a warning, so an unverified load is never silent.

Because a Hugging Face download is Hub-attested, its transport is also hardened so the attestation can't be sidestepped: a Hugging Face URL served over plaintext `http://`, or whose redirect chain downgrades to `http://`, is rejected with an `INSECURE_MODEL_SOURCE` error (loopback excepted). This only affects `huggingface.co` / `hf.co` sources — which are HTTPS in practice — and does not apply to bring-your-own HTTP. To extend the same rule to every HTTP source, set `requireSecureTransport: true` in the [configuration](/sdk/configuration#options); bring-your-own HTTP must then use HTTPS too (loopback still excepted).

### Require a verified checksum

Set `requireHttpChecksum: true` in the [configuration](/sdk/configuration#options) for a stricter Hugging Face posture: a Hugging Face URL that exposes no usable SHA-256 (for example a small non-LFS file) is then **rejected** with a `CHECKSUM_UNAVAILABLE` error instead of downloading unverified. Bring-your-own HTTP is unaffected — it still downloads unverified, with a warning.

The flag defaults to `false`. A checksum **mismatch** on a Hugging Face download always fails with a `CHECKSUM_VALIDATION_FAILED` error, regardless of this flag.

Both `requireHttpChecksum` and `requireSecureTransport` can also be set **per call** on `loadModel()` and `downloadAsset()`, overriding the config for that one download:

```ts
await loadModel({
  modelSrc: "https://huggingface.co/org/repo/resolve/main/model.gguf",
  modelType: "llamacpp-completion",
  requireSecureTransport: true,
  requireHttpChecksum: true
});
```

<Callout type="warn">
  These flags govern **downloads**. A model already in the cache is served after a freshness check without re-verification, so tightening `requireHttpChecksum` does not retroactively verify a cache entry that was fetched before the flag was set (or before this feature shipped). Clear the cache when you tighten the setting if you need the on-disk copy re-verified.
</Callout>

## Cancel a specific call by `requestId`

Both `downloadAsset()` and `loadModel()` return `Promise<string> & { requestId: string }`. The await result is unchanged (the asset path or model id, respectively), but `op.requestId` is available **synchronously** before `await` resolves — so a stop button can be wired immediately, before the first progress event arrives:

```ts
const op = downloadAsset({ assetSrc: "https://example.com/big.gguf" });
op.requestId; // synchronously available, before await

// Pause: preserves the partial file for automatic resume on the next call.
stopButton.onclick = () => cancel({ requestId: op.requestId });

// Or: discard the partial file along with the cancel.
clearButton.onclick = () => cancel({ requestId: op.requestId, clearCache: true });

await op; // rejects with InferenceCancelledError if cancelled
```

When two callers request the same artifact, the SDK deduplicates them onto a single underlying transfer. `cancel({ requestId })` rejects only the cancelling subscriber's promise; the underlying transfer keeps running to serve any other subscribers. The transfer is aborted only when the **last** subscriber leaves.

For the broader cancellation contract (errors, decorated-promise pattern across other SDK operations, broad cancel by `modelId`), see [Cancellation](/sdk/runtime/cancellation).

## Flow

* Pause: call [`cancel()`](/sdk/reference/api#cancel) with `requestId: op.requestId` from the decorated promise returned by `downloadAsset()` / `loadModel()`.
* Resume: run the same `downloadAsset()` / `loadModel()` call again — the SDK will reuse the partial file and continue downloading.
* Discard partial file: call `cancel({ requestId: op.requestId, clearCache: true })`.

## Example

The following script shows an example of pausing and resuming a download using `cancel({ requestId })` + the decorated-promise pattern:

<Tabs>
  <Tab value="js" label="JavaScript" default>
    <WrapCode>
      ```js file=<rootDir>/packages/sdk/dist/examples/download-with-cancel.js title="download-lifecycle.js" lineNumbers
      import { cancel, close, downloadAsset, LLAMA_3_2_1B_INST_Q4_0 } from '@qvac/sdk';
      console.log(`▸ Starting download with pause/resume example`);
      console.log(`\n▸ Press Ctrl+C to pause the download (it will resume on restart)\n`);
      let modelId;
      let cancelled = false;
      try {
          // Download model with progress tracking and cancellation. The
          // `downloadAsset(...)` call returns a *decorated* promise: the
          // promise resolves to the modelId, and the same value carries a
          // synchronous `requestId` field so we can cancel before it settles.
          const download = downloadAsset({
              assetSrc: LLAMA_3_2_1B_INST_Q4_0,
              onProgress: (p) => {
                  const mb = (n) => (n / 1e6).toFixed(1);
                  const line = `▸ Downloading ${p.percentage.toFixed(0)}% (${mb(p.downloaded)}/${mb(p.total)} MB)`;
                  process.stderr.write(process.stderr.isTTY ? `\r${line}` : `${line}\n`);
                  if (p.percentage >= 100)
                      process.stderr.write('\n');
                  // Example: Stops at 10% (or use Ctrl+C for manual stop)
                  if (p.percentage >= 10 && !cancelled) {
                      console.log('\n▸ Auto-cancelling at 10% for demo purposes...');
                      cancelled = true;
                      void cancel({
                          requestId: download.requestId
                          // clearCache: true, // Uncomment to delete partial file instead of resuming
                      });
                  }
              }
          });
          modelId = await download;
          console.log(`\n▸ Model downloaded successfully! Model ID: ${modelId}`);
          console.log('▸ Download completed without interruption');
          void close();
      }
      catch (error) {
          if (error instanceof Error && error.message.includes('cancelled')) {
              console.log('▸ Download was successfully cancelled');
              void close();
          }
          else {
              console.error('✖', error);
              process.exit(1);
          }
      }
      ```
    </WrapCode>
  </Tab>

  <Tab value="ts" label="TypeScript">
    <WrapCode>
      ```ts file=<rootDir>/packages/sdk/examples/download-with-cancel.ts title="download-lifecycle.ts" lineNumbers
      import { cancel, close, downloadAsset, LLAMA_3_2_1B_INST_Q4_0 } from '@qvac/sdk'

      console.log(`▸ Starting download with pause/resume example`)
      console.log(`\n▸ Press Ctrl+C to pause the download (it will resume on restart)\n`)

      let modelId: string | undefined
      let cancelled = false

      try {
        // Download model with progress tracking and cancellation. The
        // `downloadAsset(...)` call returns a *decorated* promise: the
        // promise resolves to the modelId, and the same value carries a
        // synchronous `requestId` field so we can cancel before it settles.
        const download = downloadAsset({
          assetSrc: LLAMA_3_2_1B_INST_Q4_0,
          onProgress: (p) => {
            const mb = (n: number) => (n / 1e6).toFixed(1)
            const line = `▸ Downloading ${p.percentage.toFixed(0)}% (${mb(p.downloaded)}/${mb(p.total)} MB)`
            process.stderr.write(process.stderr.isTTY ? `\r${line}` : `${line}\n`)
            if (p.percentage >= 100) process.stderr.write('\n')

            // Example: Stops at 10% (or use Ctrl+C for manual stop)
            if (p.percentage >= 10 && !cancelled) {
              console.log('\n▸ Auto-cancelling at 10% for demo purposes...')
              cancelled = true

              void cancel({
                requestId: download.requestId
                // clearCache: true, // Uncomment to delete partial file instead of resuming
              })
            }
          }
        })
        modelId = await download

        console.log(`\n▸ Model downloaded successfully! Model ID: ${modelId}`)
        console.log('▸ Download completed without interruption')
        void close()
      } catch (error) {
        if (error instanceof Error && error.message.includes('cancelled')) {
          console.log('▸ Download was successfully cancelled')
          void close()
        } else {
          console.error('✖', error)
          process.exit(1)
        }
      }
      ```
    </WrapCode>
  </Tab>

  <Tab value="python" label="Python">
    <WrapCode>
      ```python file=<rootDir>/packages/sdk-python/examples/model_info.py title="model_info.py" lineNumbers
      """Python port of packages/sdk/examples/cache-management.ts.

      Inspect a model's cache/load state, download it if missing, load, unload, and
      re-inspect. `get_model_info` and `download_asset_with_progress` are generated
      methods; `load_model`/`unload_model` are the ergonomic wrappers. Request models
      are re-exported from the flat surface.

      RUN: python examples/model_info.py
      """

      from __future__ import annotations

      import asyncio
      import sys

      from tetherto.qvac_sdk import (
          Client,
          DownloadAssetRequest,
          GetModelInfoRequest,
          ModelProgressResponse,
          download_asset_with_progress,
          get_model_info,
          load_model,
          unload_model,
      )
      from tetherto.qvac_sdk.models import WHISPER_TINY

      MB = 1024 * 1024


      def print_progress(p) -> None:
          """Print model download progress; pass as `on_progress=` to `load_model`."""
          line = (
              f"▸ Downloading {p.percentage:.0f}% "
              f"({p.downloaded / 1e6:.1f}/{p.total / 1e6:.1f} MB)"
          )
          print(line, end="\r" if sys.stderr.isatty() else "\n", file=sys.stderr)
          if p.percentage >= 100:
              print(file=sys.stderr)


      def print_status(info, label) -> None:
          print(f"\n▸ {label}")
          print(f"▸ Model Name: {info.name}")
          print(f"▸ Model ID: {info.model_id}")
          print(f"▸ Expected Size: {info.expected_size / MB:.2f} MB")
          print(f"▸ Addon: {info.addon}")
          print(f"▸ Cache Files: {len(info.cache_files or [])}")
          print(f"▸ Is Cached: {'yes' if info.is_cached else 'no'}")
          print(f"▸ Is Loaded: {'yes' if info.is_loaded else 'no'}")
          if info.is_cached and info.actual_size is not None:
              print(f"▸ Actual Size: {info.actual_size / MB:.2f} MB")


      async def fetch_info(t):
          response = await get_model_info(
              t,
              GetModelInfoRequest.model_validate(
                  {"type": "getModelInfo", "name": WHISPER_TINY.name}
              ),
          )
          return response.model_info


      async def main() -> int:
          async with Client() as client:
              t = client.transport
              try:
                  print("▸ Model Info + Cache Management Demo")

                  print("\n▸ 1. INITIAL STATUS CHECK")
                  initial = await fetch_info(t)
                  print_status(initial, "Initial Status:")

                  if not initial.is_cached:
                      print("\n▸ 2. DOWNLOADING MODEL (not cached)")
                      request = DownloadAssetRequest.model_validate(
                          {"type": "downloadAsset", "assetSrc": WHISPER_TINY.src}
                      )
                      async for event in download_asset_with_progress(t, request):
                          if isinstance(event, ModelProgressResponse):
                              print_progress(event)
                      print("▸ Download complete!")
                      print_status(await fetch_info(t), "Status After Download:")
                  else:
                      print("\n▸ 2. MODEL ALREADY CACHED — skipping download")

                  print("\n▸ 3. LOADING MODEL INTO MEMORY")
                  model_id = await load_model(t, model_src=WHISPER_TINY)
                  print(f"▸ Model loaded! ID: {model_id}")
                  print_status(await fetch_info(t), "Status After Load:")

                  print("\n▸ 4. UNLOADING MODEL")
                  await unload_model(t, model_id)
                  print("▸ Model unloaded!")
                  print_status(await fetch_info(t), "Status After Unload:")

                  print("\n▸ Demo Complete")
              except Exception as error:
                  print(f"✖ {error}", file=sys.stderr)
                  return 1
          return 0


      if __name__ == "__main__":
          sys.exit(asyncio.run(main()))
      ```
    </WrapCode>
  </Tab>
</Tabs>

<Callout type="info">
  The Python example above (`model_info.py`) covers related model-info and download-progress operations (`get_model_info`, `download_asset_with_progress`). For the exact pause/resume flow shown in the JavaScript/TypeScript tabs, use `cancel(request_id=...)` from the Python SDK — the surface is identical to `cancel()` in JS/TS.
</Callout>

<Callout type="success">
  **Tip:** all examples throughout this documentation are self-contained and runnable. For instructions on how to run them, see the [JS/TS quickstart](/sdk/js-ts-sdk#quickstart) or the [Python quickstart](/sdk/python-sdk#quickstart).
</Callout>
