---
title: "OCR"
canonical: https://docs.qvac.tether.io/sdk/v0.20/ai-capabilities/ocr/
collection: "SDK"
package: "@qvac/sdk"
line: v0.20
current_line: false
---

# OCR (/sdk/v0.20/ai-capabilities/ocr)



## Overview

OCR uses **ONNX runtime** as the inference engine. It runs a two-stage pipeline and requires compatible models for both stages:

* **Text detection**: locate text regions in an image
* **Text recognition**: decode characters in detected regions

Load supported models using `modelType: "ocr"`. Then, provide an image as either a file path (string) or an in-memory buffer. Each OCR block contains extracted text and may include `bbox` (bounding box coordinates) and `confidence` (recognition score).

## Functions

Use the following sequence of function calls:

1. [`loadModel()`](/sdk/v0.20/reference/api#loadmodel)
2. [`ocr()`](/sdk/v0.20/reference/api#ocr)
3. [`unloadModel()`](/sdk/v0.20/reference/api#unloadmodel)

For how to use each function, see [SDK — API reference](/sdk/v0.20/reference/api/).

## Models

You can load any ONNX Runtime-compatible OCR pipeline. Required files: `detector_craft.onnx` + `recognizer_<lang>.onnx` (file format: `*.onnx`).

For models available as constants, see [SDK — Models](/sdk/v0.20/#models).

## Example

The following script shows an example of OCR:

<Tabs>
  <Tab value="js" label="JavaScript" default>
    <WrapCode>
      ```js file=<rootDir>/packages/sdk/dist/examples/ocr-fasttext.js title="ocr.js" lineNumbers
      /**
       * OCR example using the QVAC SDK.
       *
       * Usage:
       *   bun examples/ocr-fasttext.ts [path-to-image]
       *
       * This example requires a test image (default: examples/image/basic_test.bmp).
       * Sample images are available in the QVAC source repository, but not included in the published npm package.
       * Pass a custom image path, or download the default image into examples/image/:
       *   https://github.com/tetherto/qvac/blob/main/packages/sdk/examples/image/basic_test.bmp
       */
      import { close, loadModel, ocr, OCR_LATIN, unloadModel } from '@qvac/sdk';
      import path from 'path';
      import { fileURLToPath } from 'url';
      const __dirname = path.dirname(fileURLToPath(import.meta.url));
      const imagePath = process.argv[2] || path.join(__dirname, 'image/basic_test.bmp');
      try {
          console.log('▸ Loading OCR model...');
          const modelId = await loadModel({
              modelSrc: OCR_LATIN,
              modelConfig: {
                  langList: ['en'],
                  magRatio: 1.5,
                  defaultRotationAngles: [90, 180, 270],
                  contrastRetry: false,
                  lowConfidenceThreshold: 0.5,
                  recognizerBatchSize: 1
              }
          });
          console.log(`▸ Model loaded successfully! Model ID: ${modelId}`);
          console.log(`\n▸ Running OCR on: ${imagePath}`);
          const { blocks } = ocr({
              modelId,
              image: imagePath,
              options: {
                  paragraph: false
              }
          });
          const result = await blocks;
          console.log('\n▸ OCR Results:');
          console.log('▸ ================================');
          for (const block of result) {
              console.log(block.text);
              if (block.bbox) {
                  console.log(`▸ BBox: [${block.bbox.join(', ')}]`);
              }
              if (block.confidence !== undefined) {
                  console.log(`▸ Confidence: ${block.confidence}`);
              }
          }
          console.log('\n▸ ================================');
          console.log('\n▸ Unloading model...');
          await unloadModel({ modelId, clearStorage: false });
          console.log('▸ Model unloaded successfully.');
          process.exit(0);
      }
      catch (error) {
          console.error('✖', error);
          await close();
      }
      ```
    </WrapCode>
  </Tab>

  <Tab value="ts" label="TypeScript">
    <WrapCode>
      ```ts file=<rootDir>/packages/sdk/examples/ocr-fasttext.ts title="ocr.ts" lineNumbers
      /**
       * OCR example using the QVAC SDK.
       *
       * Usage:
       *   bun examples/ocr-fasttext.ts [path-to-image]
       *
       * This example requires a test image (default: examples/image/basic_test.bmp).
       * Sample images are available in the QVAC source repository, but not included in the published npm package.
       * Pass a custom image path, or download the default image into examples/image/:
       *   https://github.com/tetherto/qvac/blob/main/packages/sdk/examples/image/basic_test.bmp
       */
      import { close, loadModel, ocr, OCR_LATIN, unloadModel } from '@qvac/sdk'
      import path from 'path'
      import { fileURLToPath } from 'url'

      const __dirname = path.dirname(fileURLToPath(import.meta.url))
      const imagePath = process.argv[2] || path.join(__dirname, 'image/basic_test.bmp')

      try {
        console.log('▸ Loading OCR model...')
        const modelId = await loadModel({
          modelSrc: OCR_LATIN,
          modelConfig: {
            langList: ['en'],
            magRatio: 1.5,
            defaultRotationAngles: [90, 180, 270],
            contrastRetry: false,
            lowConfidenceThreshold: 0.5,
            recognizerBatchSize: 1
          }
        })
        console.log(`▸ Model loaded successfully! Model ID: ${modelId}`)

        console.log(`\n▸ Running OCR on: ${imagePath}`)
        const { blocks } = ocr({
          modelId,
          image: imagePath,
          options: {
            paragraph: false
          }
        })

        const result = await blocks

        console.log('\n▸ OCR Results:')
        console.log('▸ ================================')
        for (const block of result) {
          console.log(block.text)
          if (block.bbox) {
            console.log(`▸ BBox: [${block.bbox.join(', ')}]`)
          }
          if (block.confidence !== undefined) {
            console.log(`▸ Confidence: ${block.confidence}`)
          }
        }
        console.log('\n▸ ================================')
        console.log('\n▸ Unloading model...')
        await unloadModel({ modelId, clearStorage: false })
        console.log('▸ Model unloaded successfully.')
        process.exit(0)
      } catch (error) {
        console.error('✖', error)
        await close()
      }
      ```
    </WrapCode>
  </Tab>

  <Tab value="python" label="Python">
    <WrapCode>
      ```python file=<rootDir>/packages/sdk-python/examples/ocr.py title="ocr.py" lineNumbers
      """Python port of packages/sdk/examples/ocr-fasttext.ts.

      Run OCR over an image and print the recognized text blocks. `ocr_stream` is a
      server-stream: it yields `OcrStreamResponse` frames, each with a `blocks` list
      whose items carry `text` (and, when present, `bbox` / `confidence`).

      Pass an image path (BMP/PNG/JPG); the SDK repo ships a sample at
      packages/sdk/examples/image/basic_test.bmp, not included in the wheel:
        python examples/ocr.py path/to/image.png
      """

      from __future__ import annotations

      import asyncio
      import sys

      from tetherto.qvac_sdk import (
          Client,
          OcrStreamRequest,
          load_model,
          ocr_stream,
          unload_model,
      )
      from tetherto.qvac_sdk.models import OCR_LATIN


      def print_progress(p) -> None:
          """Print model download progress; pass as `on_progress=` to `load_model`."""
          line = (
              f"▸ Downloading {p.percentage:.0f}% "
              f"({p.downloaded / 1e6:.1f}/{p.total / 1e6:.1f} MB)"
          )
          print(line, end="\r" if sys.stderr.isatty() else "\n", file=sys.stderr)
          if p.percentage >= 100:
              print(file=sys.stderr)


      async def main() -> int:
          if len(sys.argv) < 2:
              print("Usage: python examples/ocr.py <image-path>", file=sys.stderr)
              return 1
          image_path = sys.argv[1]

          async with Client() as client:
              t = client.transport
              try:
                  print("▸ Loading OCR model...")
                  model_id = await load_model(
                      t,
                      model_src=OCR_LATIN,
                      model_config={
                          "langList": ["en"],
                          "magRatio": 1.5,
                          "defaultRotationAngles": [90, 180, 270],
                          "contrastRetry": False,
                          "lowConfidenceThreshold": 0.5,
                          "recognizerBatchSize": 1,
                      },
                      on_progress=print_progress,
                  )
                  print(f"▸ Model loaded: {model_id}")

                  print(f"\n▸ Running OCR on: {image_path}")
                  # image is a typed chunk, not a bare path: JS's ocr wrapper turns a
                  # string into {type:"filePath"} for you, but Python calls the
                  # generated stub directly, so build the chunk explicitly.
                  request = OcrStreamRequest.model_validate(
                      {
                          "type": "ocrStream",
                          "modelId": model_id,
                          "image": {"type": "filePath", "value": image_path},
                          "options": {"paragraph": False},
                      }
                  )

                  print("\n▸ OCR Results:")
                  print("▸ ================================")
                  async for response in ocr_stream(t, request):
                      for block in response.blocks or []:
                          print(block.text)
                          bbox = getattr(block, "bbox", None)
                          if bbox is not None:
                              print(f"▸ BBox: {bbox}")
                          confidence = getattr(block, "confidence", None)
                          if confidence is not None:
                              print(f"▸ Confidence: {confidence}")
                  print("▸ ================================")

                  print("\n▸ Unloading model...")
                  await unload_model(t, model_id)
                  print("▸ Model unloaded successfully.")
              except Exception as error:
                  print(f"✖ {error}", file=sys.stderr)
                  return 1
          return 0


      if __name__ == "__main__":
          sys.exit(asyncio.run(main()))
      ```
    </WrapCode>
  </Tab>
</Tabs>

<Callout type="success">
  **Tip:** all examples throughout this documentation are self-contained and runnable. For instructions on how to run them, see the [JS/TS quickstart](/sdk/v0.20/js-ts-sdk#quickstart) or the [Python quickstart](/sdk/v0.20/python-sdk#quickstart).
</Callout>
