---
title: "Image generation"
canonical: https://docs.qvac.tether.io/sdk/ai-capabilities/image-generation/
collection: "SDK"
package: "@qvac/sdk"
line: v0.21
current_line: true
---

# Image generation (/sdk/ai-capabilities/image-generation)



## Overview

Image generation runs on a **customized Diffusion engine** ([`qvac-ext-stable-diffusion.cpp`](https://github.com/tetherto/qvac-ext-stable-diffusion.cpp)). Load a supported model using `modelType: "diffusion"`. Then, provide a text `prompt` describing the image to generate.

For image-to-image, also pass `init_image` (a `Uint8Array` of PNG or JPEG bytes) — the model transforms the input guided by the prompt instead of starting from noise. `diffusion()` returns one or more PNG images as `Uint8Array` buffers. Use `progressStream` to track generation progress step-by-step.

For higher-resolution outputs, you can chain a Real-ESRGAN upscaler. Either attach it as a post-processing step at load time and trigger it per generation via `diffusion({ upscale })`, or load an ESRGAN model standalone with `modelConfig.mode: "upscale"` and call `upscale()` on any PNG/JPEG buffer — generated or not. Both paths return PNG `Uint8Array` buffers and accept a `repeats` option that compounds the scale factor across sequential passes.

## Functions

Use the following sequence of function calls:

1. [`loadModel()`](/sdk/reference/api#loadmodel)
2. [`diffusion()`](/sdk/reference/api#diffusion) or `upscale()`
3. [`unloadModel()`](/sdk/reference/api#unloadmodel)

For how to use each function, see [SDK — API reference](/sdk/reference/api/).

## Models

Supported model families and their file layouts:

* **FLUX.2-klein**: split layout — diffusion model `*.gguf` + LLM text encoder `*.gguf` (via `llmModelSrc`) + VAE `*.safetensors` (via `vaeModelSrc`).
* **Ideogram 4**: split layout — diffusion model `*.gguf` + unconditional diffusion model `*.gguf` (via `uncondModelSrc`) + LLM text encoder `*.gguf` (via `llmModelSrc`) + VAE `*.safetensors` (via `vaeModelSrc`).
* **SD1.x, SD2.x**: single all-in-one `*.gguf` file. No companion files needed.
* **SDXL, SD3**: may require separate CLIP/T5 text encoder files (`clipLModelSrc`, `clipGModelSrc`, `t5XxlModelSrc`) in `modelConfig` depending on the model variant.
* **ESRGAN**: post-generation upscaler, `*.pth` format. Two ways to load:
  * **Paired with a diffusion model** — set `modelConfig.upscaler.model_src` at `loadModel()` time alongside the diffusion `modelSrc`.
  * **Standalone** — pass the ESRGAN model as the top-level `modelSrc` with `modelConfig.mode: "upscale"`.

For models available as constants, see [SDK — Models](/sdk#models).

<Callout type="info">
  **On upscaling:** ESRGAN can be used standalone via `modelConfig.mode: "upscale"` + `upscale()`, or optionally paired with any supported diffusion family to upscale generated images in the same call via `diffusion({ upscale })`. Available constants: `REALESRGAN_X4PLUS_ANIME_6B`, `REALESRGAN_X4PLUS`.
</Callout>

## Examples

### FLUX.2-klein

The following script shows text-to-image generation using FLUX.2-klein with its split-layout model (separate diffusion model, LLM text encoder, and VAE):

<Tabs>
  <Tab value="js" label="JavaScript" default>
    <WrapCode>
      ```js file=<rootDir>/packages/sdk/dist/examples/diffusion-flux2-klein.js title="diffusion-flux2-klein.js" lineNumbers
      import { loadModel, unloadModel, diffusion, FLUX_2_KLEIN_4B_Q4_0, FLUX_2_KLEIN_4B_VAE, QWEN3_4B_Q4_K_M } from '@qvac/sdk';
      import fs from 'fs';
      import path from 'path';
      // FLUX.2 [klein] uses a split-layout: separate diffusion model + LLM text encoder + VAE
      const diffusionModelSrc = process.argv[2] || FLUX_2_KLEIN_4B_Q4_0;
      const llmModelSrc = process.argv[3] || QWEN3_4B_Q4_K_M;
      const vaeModelSrc = process.argv[4] || FLUX_2_KLEIN_4B_VAE;
      const prompt = process.argv[5] || 'a futuristic city at sunset, photorealistic';
      const outputDir = process.argv[6] || '.';
      console.log('▸ Loading FLUX.2 [klein] split-layout model...');
      const modelId = await loadModel({
          modelSrc: diffusionModelSrc,
          modelType: 'sdcpp-generation',
          modelConfig: {
              device: 'gpu',
              threads: 4,
              llmModelSrc,
              vaeModelSrc
          },
          onProgress: (p) => {
              const mb = (n) => (n / 1e6).toFixed(1);
              const line = `▸ Downloading ${p.percentage.toFixed(0)}% (${mb(p.downloaded)}/${mb(p.total)} MB)`;
              process.stderr.write(process.stderr.isTTY ? `\r${line}` : `${line}\n`);
              if (p.percentage >= 100)
                  process.stderr.write('\n');
          }
      });
      console.log(`▸ Model loaded: ${modelId}`);
      console.log(`▸ Generating: "${prompt}"`);
      const { progressStream, outputs, stats } = diffusion({
          modelId,
          prompt,
          width: 512,
          height: 512,
          steps: 20,
          guidance: 3.5,
          cfg_scale: 1,
          seed: -1
      });
      for await (const { step, totalSteps } of progressStream) {
          console.log(`▸ step ${step}/${totalSteps}`);
      }
      const buffers = await outputs;
      for (let i = 0; i < buffers.length; i++) {
          const outputPath = path.join(outputDir, `flux2_${i}.png`);
          fs.writeFileSync(outputPath, buffers[i]);
          console.log(`▸ Saved ${outputPath}`);
      }
      console.log('▸ Stats:', await stats);
      await unloadModel({ modelId, clearStorage: false });
      console.log('▸ Done');
      process.exit(0);
      ```
    </WrapCode>
  </Tab>

  <Tab value="ts" label="TypeScript">
    <WrapCode>
      ```ts file=<rootDir>/packages/sdk/examples/diffusion-flux2-klein.ts title="diffusion-flux2-klein.ts" lineNumbers
      import {
        loadModel,
        unloadModel,
        diffusion,
        FLUX_2_KLEIN_4B_Q4_0,
        FLUX_2_KLEIN_4B_VAE,
        QWEN3_4B_Q4_K_M
      } from '@qvac/sdk'
      import fs from 'fs'
      import path from 'path'

      // FLUX.2 [klein] uses a split-layout: separate diffusion model + LLM text encoder + VAE
      const diffusionModelSrc = process.argv[2] || FLUX_2_KLEIN_4B_Q4_0
      const llmModelSrc = process.argv[3] || QWEN3_4B_Q4_K_M
      const vaeModelSrc = process.argv[4] || FLUX_2_KLEIN_4B_VAE
      const prompt = process.argv[5] || 'a futuristic city at sunset, photorealistic'
      const outputDir = process.argv[6] || '.'

      console.log('▸ Loading FLUX.2 [klein] split-layout model...')

      const modelId = await loadModel({
        modelSrc: diffusionModelSrc,
        modelType: 'sdcpp-generation',
        modelConfig: {
          device: 'gpu',
          threads: 4,
          llmModelSrc,
          vaeModelSrc
        },
        onProgress: (p) => {
          const mb = (n: number) => (n / 1e6).toFixed(1)
          const line = `▸ Downloading ${p.percentage.toFixed(0)}% (${mb(p.downloaded)}/${mb(p.total)} MB)`
          process.stderr.write(process.stderr.isTTY ? `\r${line}` : `${line}\n`)
          if (p.percentage >= 100) process.stderr.write('\n')
        }
      })
      console.log(`▸ Model loaded: ${modelId}`)

      console.log(`▸ Generating: "${prompt}"`)

      const { progressStream, outputs, stats } = diffusion({
        modelId,
        prompt,
        width: 512,
        height: 512,
        steps: 20,
        guidance: 3.5,
        cfg_scale: 1,
        seed: -1
      })

      for await (const { step, totalSteps } of progressStream) {
        console.log(`▸ step ${step}/${totalSteps}`)
      }

      const buffers = await outputs
      for (let i = 0; i < buffers.length; i++) {
        const outputPath = path.join(outputDir, `flux2_${i}.png`)
        fs.writeFileSync(outputPath, buffers[i]!)
        console.log(`▸ Saved ${outputPath}`)
      }

      console.log('▸ Stats:', await stats)
      await unloadModel({ modelId, clearStorage: false })
      console.log('▸ Done')
      process.exit(0)
      ```
    </WrapCode>
  </Tab>
</Tabs>

### Ideogram 4

Ideogram 4 uses a split layout with an extra **unconditional diffusion model** (`uncondModelSrc`) alongside the Qwen3-VL LLM text encoder (`llmModelSrc`) and VAE (`vaeModelSrc`). Unlike the other families, its `prompt` must be a **JSON-serialized structured caption** — an object with a `high_level_description`, a `style_description`, and a `compositional_deconstruction` that lays out each element with an explicit bounding box (`bbox`) and, for text elements, the literal `text` to render. Plain-text prompts produce degenerate or placeholder output.

<Callout type="warn">
  Pass the structured caption through `JSON.stringify(...)` when calling `diffusion({ prompt })`. The bounding boxes are expressed in the target canvas coordinate space (`[x0, y0, x1, y1]`), so keep them consistent with the requested `width`/`height`. Failure to abide by the JSON format specified within the examples, or adding NSFW prompts, will result in censored image output.
</Callout>

The following script loads Ideogram 4 in split-layout (diffusion model + unconditional model + LLM text encoder + VAE) and generates from a structured caption:

<Tabs>
  <Tab value="js" label="JavaScript" default>
    <WrapCode>
      ```js file=<rootDir>/packages/sdk/dist/examples/diffusion-ideogram.js title="diffusion-ideogram.js" lineNumbers
      import { diffusion, loadModel, unloadModel } from '@qvac/sdk';
      import fs from 'fs';
      import path from 'path';
      const diffusionModelSrc = process.argv[2] ||
          'https://huggingface.co/leejet/ideogram-4-GGUF/resolve/main/ideogram4-Q4_0.gguf';
      const uncondModelSrc = process.argv[3] ||
          'https://huggingface.co/leejet/ideogram-4-GGUF/resolve/main/ideogram4_uncond-Q4_0.gguf';
      const llmModelSrc = process.argv[4] ||
          'https://huggingface.co/unsloth/Qwen3-VL-8B-Instruct-GGUF/resolve/main/Qwen3-VL-8B-Instruct-Q4_K_M.gguf';
      const vaeModelSrc = process.argv[5] ||
          'https://huggingface.co/Comfy-Org/flux2-dev/resolve/main/split_files/vae/flux2-vae.safetensors';
      const outputDir = process.argv[6] || '.';
      function formatMb(bytes) {
          return (bytes / 1e6).toFixed(1);
      }
      // Ideogram 4 expects a JSON-serialized structured caption with explicit
      // bounding boxes. Plain-text prompts produce degenerate or placeholder output.
      const structuredPrompt = {
          high_level_description: 'A bright product photo of one yellow coffee mug on a clean office desk with a tiny readable label that says IDEO.',
          style_description: {
              aesthetics: 'simple product photography, clean bright office desk, minimal composition',
              lighting: 'soft studio light with gentle shadows',
              photo: 'sharp focus, crisp readable label text',
              medium: 'product photo',
              color_palette: ['#FFD54A', '#FFFFFF', '#111111', '#7EC8E3']
          },
          compositional_deconstruction: {
              canvas: 'Square canvas, upright orientation. All text is horizontal and readable.',
              background: 'Clean white office desk with a soft blue background gradient.',
              elements: [
                  {
                      type: 'obj',
                      bbox: [230, 260, 800, 740],
                      desc: 'Exactly one matte yellow ceramic coffee mug centered on the desk.'
                  },
                  {
                      type: 'text',
                      bbox: [510, 360, 620, 640],
                      text: 'IDEO',
                      desc: 'Small black label printed horizontally on the mug.'
                  }
              ]
          }
      };
      if (process.argv.includes('--help')) {
          console.log('Usage: tsx examples/diffusion-ideogram.ts [model] [uncond] [llm] [vae] [output-dir]');
          console.log('▸ Defaults use the four Hugging Face model URLs supported by the diffusion addon.');
          process.exit(0);
      }
      let modelId;
      try {
          console.log('▸ Loading Ideogram 4 split-layout model...');
          modelId = await loadModel({
              modelSrc: diffusionModelSrc,
              modelType: 'sdcpp-generation',
              modelConfig: {
                  device: 'gpu',
                  threads: 4,
                  diffusion_fa: true,
                  offload_to_cpu: true,
                  llmModelSrc,
                  vaeModelSrc,
                  uncondModelSrc
              },
              onProgress: (progress) => {
                  const line = `▸ Downloading ${progress.percentage.toFixed(0)}% (${formatMb(progress.downloaded)}/${formatMb(progress.total)} MB)`;
                  process.stderr.write(process.stderr.isTTY ? `\r${line}` : `${line}\n`);
                  if (progress.percentage >= 100)
                      process.stderr.write('\n');
              }
          });
          console.log(`▸ Model loaded: ${modelId}`);
          const { progressStream, outputs, stats } = diffusion({
              modelId,
              prompt: JSON.stringify(structuredPrompt),
              width: 768,
              height: 768,
              steps: 16,
              cfg_scale: 7,
              seed: 42
          });
          for await (const { step, totalSteps } of progressStream) {
              console.log(`▸ step ${step}/${totalSteps}`);
          }
          const buffers = await outputs;
          for (let i = 0; i < buffers.length; i++) {
              const outputPath = path.join(outputDir, `ideogram_${i}.png`);
              fs.writeFileSync(outputPath, buffers[i]);
              console.log(`▸ Saved ${outputPath}`);
          }
          console.log('▸ Stats:', await stats);
          await unloadModel({ modelId, clearStorage: false });
          modelId = undefined;
          console.log('▸ Done');
          process.exit(0);
      }
      catch (error) {
          if (modelId) {
              try {
                  await unloadModel({ modelId, clearStorage: false });
              }
              catch {
                  // Preserve the original error.
              }
          }
          console.error('✖', error);
          process.exit(1);
      }
      ```
    </WrapCode>
  </Tab>

  <Tab value="ts" label="TypeScript">
    <WrapCode>
      ```ts file=<rootDir>/packages/sdk/examples/diffusion-ideogram.ts title="diffusion-ideogram.ts" lineNumbers
      import { diffusion, loadModel, unloadModel } from '@qvac/sdk'
      import fs from 'fs'
      import path from 'path'

      const diffusionModelSrc =
        process.argv[2] ||
        'https://huggingface.co/leejet/ideogram-4-GGUF/resolve/main/ideogram4-Q4_0.gguf'
      const uncondModelSrc =
        process.argv[3] ||
        'https://huggingface.co/leejet/ideogram-4-GGUF/resolve/main/ideogram4_uncond-Q4_0.gguf'
      const llmModelSrc =
        process.argv[4] ||
        'https://huggingface.co/unsloth/Qwen3-VL-8B-Instruct-GGUF/resolve/main/Qwen3-VL-8B-Instruct-Q4_K_M.gguf'
      const vaeModelSrc =
        process.argv[5] ||
        'https://huggingface.co/Comfy-Org/flux2-dev/resolve/main/split_files/vae/flux2-vae.safetensors'
      const outputDir = process.argv[6] || '.'

      function formatMb(bytes: number) {
        return (bytes / 1e6).toFixed(1)
      }

      // Ideogram 4 expects a JSON-serialized structured caption with explicit
      // bounding boxes. Plain-text prompts produce degenerate or placeholder output.
      const structuredPrompt = {
        high_level_description:
          'A bright product photo of one yellow coffee mug on a clean office desk with a tiny readable label that says IDEO.',
        style_description: {
          aesthetics: 'simple product photography, clean bright office desk, minimal composition',
          lighting: 'soft studio light with gentle shadows',
          photo: 'sharp focus, crisp readable label text',
          medium: 'product photo',
          color_palette: ['#FFD54A', '#FFFFFF', '#111111', '#7EC8E3']
        },
        compositional_deconstruction: {
          canvas: 'Square canvas, upright orientation. All text is horizontal and readable.',
          background: 'Clean white office desk with a soft blue background gradient.',
          elements: [
            {
              type: 'obj',
              bbox: [230, 260, 800, 740],
              desc: 'Exactly one matte yellow ceramic coffee mug centered on the desk.'
            },
            {
              type: 'text',
              bbox: [510, 360, 620, 640],
              text: 'IDEO',
              desc: 'Small black label printed horizontally on the mug.'
            }
          ]
        }
      }

      if (process.argv.includes('--help')) {
        console.log('Usage: tsx examples/diffusion-ideogram.ts [model] [uncond] [llm] [vae] [output-dir]')
        console.log('▸ Defaults use the four Hugging Face model URLs supported by the diffusion addon.')
        process.exit(0)
      }

      let modelId: string | undefined

      try {
        console.log('▸ Loading Ideogram 4 split-layout model...')
        modelId = await loadModel({
          modelSrc: diffusionModelSrc,
          modelType: 'sdcpp-generation',
          modelConfig: {
            device: 'gpu',
            threads: 4,
            diffusion_fa: true,
            offload_to_cpu: true,
            llmModelSrc,
            vaeModelSrc,
            uncondModelSrc
          },
          onProgress: (progress) => {
            const line = `▸ Downloading ${progress.percentage.toFixed(0)}% (${formatMb(progress.downloaded)}/${formatMb(progress.total)} MB)`
            process.stderr.write(process.stderr.isTTY ? `\r${line}` : `${line}\n`)
            if (progress.percentage >= 100) process.stderr.write('\n')
          }
        })
        console.log(`▸ Model loaded: ${modelId}`)

        const { progressStream, outputs, stats } = diffusion({
          modelId,
          prompt: JSON.stringify(structuredPrompt),
          width: 768,
          height: 768,
          steps: 16,
          cfg_scale: 7,
          seed: 42
        })

        for await (const { step, totalSteps } of progressStream) {
          console.log(`▸ step ${step}/${totalSteps}`)
        }

        const buffers = await outputs
        for (let i = 0; i < buffers.length; i++) {
          const outputPath = path.join(outputDir, `ideogram_${i}.png`)
          fs.writeFileSync(outputPath, buffers[i]!)
          console.log(`▸ Saved ${outputPath}`)
        }

        console.log('▸ Stats:', await stats)
        await unloadModel({ modelId, clearStorage: false })
        modelId = undefined
        console.log('▸ Done')
        process.exit(0)
      } catch (error) {
        if (modelId) {
          try {
            await unloadModel({ modelId, clearStorage: false })
          } catch {
            // Preserve the original error.
          }
        }
        console.error('✖', error)
        process.exit(1)
      }
      ```
    </WrapCode>
  </Tab>
</Tabs>

### Stable Diffusion

The following script shows a minimal text-to-image generation example using a single all-in-one SD 2.1 model:

<Tabs>
  <Tab value="js" label="JavaScript" default>
    <WrapCode>
      ```js file=<rootDir>/packages/sdk/dist/examples/diffusion-simple.js title="diffusion-simple.js" lineNumbers
      import { loadModel, unloadModel, diffusion, SD_V2_1_1B_Q8_0 } from '@qvac/sdk';
      import fs from 'fs';
      // Minimal diffusion example — single GGUF model, no companion files needed.
      // Works with SD 1.x / 2.x all-in-one models.
      const modelSrc = process.argv[2] || SD_V2_1_1B_Q8_0;
      const prompt = process.argv[3] || 'a photo of a cat sitting on a windowsill';
      const modelId = await loadModel({
          modelSrc,
          modelType: 'sdcpp-generation',
          modelConfig: { prediction: 'v' }
      });
      const { outputs } = diffusion({ modelId, prompt });
      const buffers = await outputs;
      fs.writeFileSync('output.png', buffers[0]);
      console.log('▸ Saved output.png');
      await unloadModel({ modelId, clearStorage: false });
      process.exit(0);
      ```
    </WrapCode>
  </Tab>

  <Tab value="ts" label="TypeScript">
    <WrapCode>
      ```ts file=<rootDir>/packages/sdk/examples/diffusion-simple.ts title="diffusion-simple.ts" lineNumbers
      import { loadModel, unloadModel, diffusion, SD_V2_1_1B_Q8_0 } from '@qvac/sdk'
      import fs from 'fs'

      // Minimal diffusion example — single GGUF model, no companion files needed.
      // Works with SD 1.x / 2.x all-in-one models.
      const modelSrc = process.argv[2] || SD_V2_1_1B_Q8_0
      const prompt = process.argv[3] || 'a photo of a cat sitting on a windowsill'

      const modelId = await loadModel({
        modelSrc,
        modelType: 'sdcpp-generation',
        modelConfig: { prediction: 'v' }
      })

      const { outputs } = diffusion({ modelId, prompt })
      const buffers = await outputs

      fs.writeFileSync('output.png', buffers[0]!)
      console.log('▸ Saved output.png')

      await unloadModel({ modelId, clearStorage: false })
      process.exit(0)
      ```
    </WrapCode>
  </Tab>
</Tabs>

### Image-to-image

Pass `init_image` to transform an existing image guided by a text prompt. Behavior depends on the model family:

* **FLUX.2**: in-context conditioning. Requires `prediction: "flux2_flow"` in `modelConfig` at `loadModel()` time; `strength` is ignored on this path.
* **SD / SDXL / SD3**: SDEdit-style. Use `strength` to control how much the source is preserved (`0` = keep source, `1` = ignore source).

The following script loads FLUX.2-klein in split-layout and transforms an input image using in-context conditioning (`prediction: "flux2_flow"`):

<Tabs>
  <Tab value="js" label="JavaScript" default>
    <WrapCode>
      ```js file=<rootDir>/packages/sdk/dist/examples/diffusion-flux2-klein-img2img.js title="diffusion-flux2-klein-img2img.js" lineNumbers
      import { loadModel, unloadModel, diffusion, FLUX_2_KLEIN_4B_Q4_0, FLUX_2_KLEIN_4B_VAE, QWEN3_4B_Q4_K_M } from '@qvac/sdk';
      import fs from 'fs';
      import path from 'path';
      // img2img with FLUX.2 [klein] split-layout — uses in-context conditioning ("flux2_flow").
      const inputPath = process.argv[2];
      const prompt = process.argv[3] || 'oil painting style, vibrant colors';
      const outputDir = process.argv[4] || '.';
      const diffusionModelSrc = process.argv[5] || FLUX_2_KLEIN_4B_Q4_0;
      const llmModelSrc = process.argv[6] || QWEN3_4B_Q4_K_M;
      const vaeModelSrc = process.argv[7] || FLUX_2_KLEIN_4B_VAE;
      if (!inputPath) {
          console.error('✖ input image path is required');
          console.error('Usage: bun run bare:example dist/examples/diffusion-flux2-klein-img2img.js <inputImage> [prompt] [outputDir] [diffusionModelSrc] [llmModelSrc] [vaeModelSrc]');
          process.exit(1);
      }
      try {
          console.log('▸ Loading FLUX.2 [klein] split-layout model...');
          const modelId = await loadModel({
              modelSrc: diffusionModelSrc,
              modelType: 'sdcpp-generation',
              modelConfig: {
                  device: 'gpu',
                  threads: 4,
                  llmModelSrc,
                  vaeModelSrc,
                  prediction: 'flux2_flow'
              },
              onProgress: (p) => {
                  const mb = (n) => (n / 1e6).toFixed(1);
                  const line = `▸ Downloading ${p.percentage.toFixed(0)}% (${mb(p.downloaded)}/${mb(p.total)} MB)`;
                  process.stderr.write(process.stderr.isTTY ? `\r${line}` : `${line}\n`);
                  if (p.percentage >= 100)
                      process.stderr.write('\n');
              }
          });
          console.log(`▸ Model loaded: ${modelId}`);
          const init_image = new Uint8Array(fs.readFileSync(inputPath));
          console.log(`▸ Transforming "${inputPath}" with prompt: "${prompt}"`);
          const { progressStream, outputs, stats } = diffusion({
              modelId,
              prompt,
              init_image,
              steps: 20,
              guidance: 3.5,
              cfg_scale: 1,
              seed: -1
          });
          for await (const { step, totalSteps } of progressStream) {
              console.log(`▸ step ${step}/${totalSteps}`);
          }
          const buffers = await outputs;
          for (let i = 0; i < buffers.length; i++) {
              const outputPath = path.join(outputDir, `flux2_img2img_${i}.png`);
              fs.writeFileSync(outputPath, buffers[i]);
              console.log(`▸ Saved ${outputPath}`);
          }
          console.log('▸ Stats:', await stats);
          await unloadModel({ modelId, clearStorage: false });
          console.log('▸ Done');
          process.exit(0);
      }
      catch (error) {
          console.error('✖', error);
          process.exit(1);
      }
      ```
    </WrapCode>
  </Tab>

  <Tab value="ts" label="TypeScript">
    <WrapCode>
      ```ts file=<rootDir>/packages/sdk/examples/diffusion-flux2-klein-img2img.ts title="diffusion-flux2-klein-img2img.ts" lineNumbers
      import {
        loadModel,
        unloadModel,
        diffusion,
        FLUX_2_KLEIN_4B_Q4_0,
        FLUX_2_KLEIN_4B_VAE,
        QWEN3_4B_Q4_K_M
      } from '@qvac/sdk'
      import fs from 'fs'
      import path from 'path'

      // img2img with FLUX.2 [klein] split-layout — uses in-context conditioning ("flux2_flow").

      const inputPath = process.argv[2]
      const prompt = process.argv[3] || 'oil painting style, vibrant colors'
      const outputDir = process.argv[4] || '.'
      const diffusionModelSrc = process.argv[5] || FLUX_2_KLEIN_4B_Q4_0
      const llmModelSrc = process.argv[6] || QWEN3_4B_Q4_K_M
      const vaeModelSrc = process.argv[7] || FLUX_2_KLEIN_4B_VAE

      if (!inputPath) {
        console.error('✖ input image path is required')
        console.error(
          'Usage: bun run bare:example dist/examples/diffusion-flux2-klein-img2img.js <inputImage> [prompt] [outputDir] [diffusionModelSrc] [llmModelSrc] [vaeModelSrc]'
        )
        process.exit(1)
      }

      try {
        console.log('▸ Loading FLUX.2 [klein] split-layout model...')
        const modelId = await loadModel({
          modelSrc: diffusionModelSrc,
          modelType: 'sdcpp-generation',
          modelConfig: {
            device: 'gpu',
            threads: 4,
            llmModelSrc,
            vaeModelSrc,
            prediction: 'flux2_flow'
          },
          onProgress: (p) => {
            const mb = (n: number) => (n / 1e6).toFixed(1)
            const line = `▸ Downloading ${p.percentage.toFixed(0)}% (${mb(p.downloaded)}/${mb(p.total)} MB)`
            process.stderr.write(process.stderr.isTTY ? `\r${line}` : `${line}\n`)
            if (p.percentage >= 100) process.stderr.write('\n')
          }
        })
        console.log(`▸ Model loaded: ${modelId}`)

        const init_image = new Uint8Array(fs.readFileSync(inputPath))
        console.log(`▸ Transforming "${inputPath}" with prompt: "${prompt}"`)

        const { progressStream, outputs, stats } = diffusion({
          modelId,
          prompt,
          init_image,
          steps: 20,
          guidance: 3.5,
          cfg_scale: 1,
          seed: -1
        })

        for await (const { step, totalSteps } of progressStream) {
          console.log(`▸ step ${step}/${totalSteps}`)
        }

        const buffers = await outputs
        for (let i = 0; i < buffers.length; i++) {
          const outputPath = path.join(outputDir, `flux2_img2img_${i}.png`)
          fs.writeFileSync(outputPath, buffers[i]!)
          console.log(`▸ Saved ${outputPath}`)
        }

        console.log('▸ Stats:', await stats)
        await unloadModel({ modelId, clearStorage: false })
        console.log('▸ Done')
        process.exit(0)
      } catch (error) {
        console.error('✖', error)
        process.exit(1)
      }
      ```
    </WrapCode>
  </Tab>
</Tabs>

### Upscaling

#### Post-processing

Pass `upscale` to `diffusion()` to upscale generated images in the same call. Requires `modelConfig.upscaler = { type: "esrgan", model_src, tile_size? }` at `loadModel()` time. Behavior depends on the value passed:

* `true` (or `{}` / `{ repeats: 1 }`): single pass at the model's native scale factor (e.g. `x4` for RealESRGAN\_x4plus).
* `{ repeats: N }`: N sequential passes — each pass multiplies the output dimensions by the model's scale factor (e.g. `repeats: 2` with an `x4` model → `x16`).
* When `batch_count > 1`, every output image is upscaled independently.

The following script loads SD 2.1 with an ESRGAN upscaler and generates both a single-pass (`x4`) and a two-pass (`x16`) upscale from the same prompt:

<Tabs>
  <Tab value="js" label="JavaScript" default>
    <WrapCode>
      ```js file=<rootDir>/packages/sdk/dist/examples/diffusion-esrgan-upscale.js title="diffusion-esrgan-upscale.js" lineNumbers
      import { loadModel, unloadModel, diffusion, SD_V2_1_1B_Q8_0, REALESRGAN_X4PLUS_ANIME_6B } from '@qvac/sdk';
      import fs from 'fs';
      import path from 'path';
      // ESRGAN upscale example.
      //
      // Usage:
      //   bun run examples/diffusion-esrgan-upscale.ts [esrganSrc] [prompt] [outputDir]
      const esrganArg = process.argv[2];
      const promptArg = process.argv[3];
      const outputDirArg = process.argv[4];
      const esrganModelSrc = esrganArg ?? REALESRGAN_X4PLUS_ANIME_6B;
      const prompt = promptArg ??
          'an illustrated red fox portrait, clean line art, soft watercolor background, detailed fur, crisp eyes';
      const negative_prompt = 'blurry, low quality, watermark, text';
      const outputDir = outputDirArg ?? '.';
      const seed = 42;
      try {
          console.log('▸ Loading SD 2.1 + ESRGAN upscaler...');
          const modelId = await loadModel({
              modelSrc: SD_V2_1_1B_Q8_0,
              modelConfig: {
                  prediction: 'v',
                  upscaler: {
                      type: 'esrgan',
                      model_src: esrganModelSrc,
                      tile_size: 128
                  }
              },
              onProgress: (p) => {
                  const mb = (n) => (n / 1e6).toFixed(1);
                  const line = `▸ Downloading ${p.percentage.toFixed(0)}% (${mb(p.downloaded)}/${mb(p.total)} MB)`;
                  process.stderr.write(process.stderr.isTTY ? `\r${line}` : `${line}\n`);
                  if (p.percentage >= 100)
                      process.stderr.write('\n');
              }
          });
          console.log(`▸ Model loaded: ${modelId}`);
          // Source size is intentionally small — each ESRGAN repeat multiplies dimensions.
          const baseParams = {
              modelId,
              prompt,
              negative_prompt,
              width: 128,
              height: 128,
              steps: 5,
              cfg_scale: 7.5,
              seed
          };
          console.log(`▸ Generating ESRGAN x4 upscale: "${prompt}"`);
          const single = diffusion({ ...baseParams, upscale: true });
          for await (const { step, totalSteps } of single.progressStream) {
              console.log(`▸ step ${step}/${totalSteps}`);
          }
          const singleBuffers = await single.outputs;
          for (let i = 0; i < singleBuffers.length; i++) {
              const out = path.join(outputDir, `sd2_esrgan_x4_seed${seed}_${i}.png`);
              fs.writeFileSync(out, singleBuffers[i]);
              console.log(`▸ Saved ${out}`);
          }
          console.log('▸ Stats:', await single.stats);
          console.log('▸ Generating ESRGAN two-pass x16 upscale...');
          const twoPass = diffusion({ ...baseParams, upscale: { repeats: 2 } });
          for await (const { step, totalSteps } of twoPass.progressStream) {
              console.log(`▸ step ${step}/${totalSteps}`);
          }
          const twoPassBuffers = await twoPass.outputs;
          for (let i = 0; i < twoPassBuffers.length; i++) {
              const out = path.join(outputDir, `sd2_esrgan_x16_seed${seed}_${i}.png`);
              fs.writeFileSync(out, twoPassBuffers[i]);
              console.log(`▸ Saved ${out}`);
          }
          console.log('▸ Stats:', await twoPass.stats);
          await unloadModel({ modelId, clearStorage: false });
          console.log('▸ Done');
          process.exit(0);
      }
      catch (error) {
          console.error('✖', error);
          process.exit(1);
      }
      ```
    </WrapCode>
  </Tab>

  <Tab value="ts" label="TypeScript">
    <WrapCode>
      ```ts file=<rootDir>/packages/sdk/examples/diffusion-esrgan-upscale.ts title="diffusion-esrgan-upscale.ts" lineNumbers
      import {
        loadModel,
        unloadModel,
        diffusion,
        SD_V2_1_1B_Q8_0,
        REALESRGAN_X4PLUS_ANIME_6B
      } from '@qvac/sdk'
      import fs from 'fs'
      import path from 'path'

      // ESRGAN upscale example.
      //
      // Usage:
      //   bun run examples/diffusion-esrgan-upscale.ts [esrganSrc] [prompt] [outputDir]

      const esrganArg: string | undefined = process.argv[2]
      const promptArg: string | undefined = process.argv[3]
      const outputDirArg: string | undefined = process.argv[4]

      const esrganModelSrc = esrganArg ?? REALESRGAN_X4PLUS_ANIME_6B

      const prompt =
        promptArg ??
        'an illustrated red fox portrait, clean line art, soft watercolor background, detailed fur, crisp eyes'
      const negative_prompt = 'blurry, low quality, watermark, text'
      const outputDir = outputDirArg ?? '.'
      const seed = 42

      try {
        console.log('▸ Loading SD 2.1 + ESRGAN upscaler...')
        const modelId = await loadModel({
          modelSrc: SD_V2_1_1B_Q8_0,
          modelConfig: {
            prediction: 'v',
            upscaler: {
              type: 'esrgan',
              model_src: esrganModelSrc,
              tile_size: 128
            }
          },
          onProgress: (p) => {
            const mb = (n: number) => (n / 1e6).toFixed(1)
            const line = `▸ Downloading ${p.percentage.toFixed(0)}% (${mb(p.downloaded)}/${mb(p.total)} MB)`
            process.stderr.write(process.stderr.isTTY ? `\r${line}` : `${line}\n`)
            if (p.percentage >= 100) process.stderr.write('\n')
          }
        })
        console.log(`▸ Model loaded: ${modelId}`)

        // Source size is intentionally small — each ESRGAN repeat multiplies dimensions.
        const baseParams = {
          modelId,
          prompt,
          negative_prompt,
          width: 128,
          height: 128,
          steps: 5,
          cfg_scale: 7.5,
          seed
        }

        console.log(`▸ Generating ESRGAN x4 upscale: "${prompt}"`)
        const single = diffusion({ ...baseParams, upscale: true })
        for await (const { step, totalSteps } of single.progressStream) {
          console.log(`▸ step ${step}/${totalSteps}`)
        }

        const singleBuffers = await single.outputs
        for (let i = 0; i < singleBuffers.length; i++) {
          const out = path.join(outputDir, `sd2_esrgan_x4_seed${seed}_${i}.png`)
          fs.writeFileSync(out, singleBuffers[i]!)
          console.log(`▸ Saved ${out}`)
        }
        console.log('▸ Stats:', await single.stats)

        console.log('▸ Generating ESRGAN two-pass x16 upscale...')
        const twoPass = diffusion({ ...baseParams, upscale: { repeats: 2 } })
        for await (const { step, totalSteps } of twoPass.progressStream) {
          console.log(`▸ step ${step}/${totalSteps}`)
        }

        const twoPassBuffers = await twoPass.outputs
        for (let i = 0; i < twoPassBuffers.length; i++) {
          const out = path.join(outputDir, `sd2_esrgan_x16_seed${seed}_${i}.png`)
          fs.writeFileSync(out, twoPassBuffers[i]!)
          console.log(`▸ Saved ${out}`)
        }
        console.log('▸ Stats:', await twoPass.stats)

        await unloadModel({ modelId, clearStorage: false })
        console.log('▸ Done')
        process.exit(0)
      } catch (error) {
        console.error('✖', error)
        process.exit(1)
      }
      ```
    </WrapCode>
  </Tab>
</Tabs>

#### Standalone

The SDK also exposes `upscale()` — a standalone path that upscales any PNG/JPEG image without running diffusion. Load the ESRGAN model directly with `modelConfig.mode: "upscale"`:

```ts
const modelId = await loadModel(REALESRGAN_X4PLUS_ANIME_6B, {
  modelType: "diffusion",
  modelConfig: { mode: "upscale", upscaler: { tile_size: 128 } },
});
const { outputs } = upscale({ modelId, image: pngBytes, repeats: 2 });
const [upscaledPng] = await outputs;
```

`tile_size` defaults to 128. Increase it to reduce tile seams on large inputs, at the cost of more memory per pass.

<Callout type="info">
  The Python client supports this capability through the same worker. A dedicated Python example is not yet published — see the [Python SDK](/sdk/python-sdk) for the API surface.
</Callout>

<Callout type="success">
  **Tip:** all examples throughout this documentation are self-contained and runnable. For instructions on how to run them, see the [JS/TS quickstart](/sdk/js-ts-sdk#quickstart) or the [Python quickstart](/sdk/python-sdk#quickstart).
</Callout>
