---
title: "RAG"
canonical: https://docs.qvac.tether.io/sdk/v0.18/ai-capabilities/rag/
collection: "SDK"
package: "@qvac/sdk"
line: v0.18
current_line: false
---

# RAG (/sdk/v0.18/ai-capabilities/rag)



## Overview

RAG (retrieval-augmented generation) uses [text embeddings](/sdk/v0.18/ai-capabilities/text-embeddings): you embed (vectorize) documents, persist them in a vector store, and later retrieve the most relevant chunks for a query using similarity search.

Compared to generating text embeddings only, the key differences are:

* You must **persist embeddings** (vectors) alongside the original text (and optional metadata) in a vector store.
* At query time, you **embed the query** and run **top‑K** vector search to fetch the most relevant documents/chunks.
* You *typically* pass the retrieved text to [`completion()`](/sdk/v0.18/reference/api#completion) as context to ground the model's answer (this retrieval step is what makes it RAG).

## Functions

1. [`loadModel()`](/sdk/v0.18/reference/api#loadmodel) (with `modelType: "embeddings"`)
2. [`embed()`](/sdk/v0.18/reference/api#embed) — generate vectors
3. Use any combination of the RAG functions below as needed:
   * [`ragChunk()`](/sdk/v0.18/reference/api#ragchunk) — chunk documents
   * [`ragIngest()`](/sdk/v0.18/reference/api#ragingest) — embed and store (requires model)
   * [`ragSaveEmbeddings()`](/sdk/v0.18/reference/api#ragsaveembeddings) — save pre-computed vectors
   * [`ragSearch()`](/sdk/v0.18/reference/api#ragsearch) — query similar documents (requires model)
   * [`ragReindex()`](/sdk/v0.18/reference/api#ragreindex) — optimize search index
   * [`ragDeleteEmbeddings()`](/sdk/v0.18/reference/api#ragdeleteembeddings) — remove documents
   * [`ragListWorkspaces()`](/sdk/v0.18/reference/api#raglistworkspaces) — list workspaces
   * [`ragCloseWorkspace()`](/sdk/v0.18/reference/api#ragcloseworkspace) — release resources
   * [`ragDeleteWorkspace()`](/sdk/v0.18/reference/api#ragdeleteworkspace) — delete workspace and data
   * [`createVectorIndex()`](/sdk/v0.18/reference/api#createvectorindex) — build a TurboVec index over your own documents
   * [`loadVectorIndex()`](/sdk/v0.18/reference/api#loadvectorindex) — reopen a saved TurboVec index
4. [`unloadModel()`](/sdk/v0.18/reference/api#unloadmodel)

For how to use each function, see [SDK — API reference](/sdk/v0.18/reference/api/).

## Pipeline

Create your RAG pipeline using [text embeddings](/sdk/v0.18/ai-capabilities/text-embeddings), the functions above, and [completion](/sdk/v0.18/ai-capabilities/text-generation).

Regarding vector storage, you may use:

* **External vector DB:** use [`embed()`](/sdk/v0.18/reference/api#embed) to generate vectors and store them wherever you want (e.g., MongoDB, LanceDB, ChromaDB, SQLite-Vector). Choose this when you need custom persistence, filtering, or integration with an existing database.
* **Built-in TurboVec index with your own document store:** use [`embed()`](/sdk/v0.18/reference/api#embed) for vectors and [`createVectorIndex()`](/sdk/v0.18/reference/api#createvectorindex) for on-device nearest-neighbour search. QVAC indexes the vectors; your documents stay in whatever store you already have, and results come back as the ids you assigned. Snapshots can be saved and reopened with [`loadVectorIndex()`](/sdk/v0.18/reference/api#loadvectorindex).
* **Built-in vector store:** use the RAG workspace functions ([`ragIngest()`](/sdk/v0.18/reference/api#ragingest), [`ragSearch()`](/sdk/v0.18/reference/api#ragsearch), etc.). QVAC persists the vectors for you, so this is the simplest path.

<Callout type="warn">
  **Built-in vector store is not production grade** and is intended for prototypes only. For production workloads, use one of the external vector DB paths shown below (MongoDB, SQLite, LanceDB, Chroma, etc.).
</Callout>

<Callout type="info">
  **Important:** when using an external vector DB, make sure its schema matches the embedding dimensionality produced by your model (e.g., GTE Large embeddings used in the example are 1024‑dimensional).
</Callout>

## Examples

### Built-in vector store

<Callout type="warn">
  **Prototype only** — for production, prefer an external vector DB (below).
</Callout>

The following script shows how to ingest documents into a built-in RAG workspace with `ragIngest()` and query them with `ragSearch()`, without setting up an external vector DB:

<Tabs>
  <Tab value="js" label="JavaScript" default>
    <WrapCode>
      ```js file=<rootDir>/packages/sdk/dist/examples/rag/rag-hyperdb/ingest.js title="rag-managed.js" lineNumbers
      import { loadModel, unloadModel, GTE_LARGE_FP16, ragIngest, ragSearch, ragCloseWorkspace } from '@qvac/sdk';
      try {
          // Get query from command line or use default
          const query = process.argv[2] || 'machine learning algorithms';
          const workspace = 'ingest-example';
          console.log(`▸ Query: "${query}"`);
          const modelId = await loadModel({
              modelSrc: GTE_LARGE_FP16,
              onProgress: (p) => {
                  const mb = (n) => (n / 1e6).toFixed(1);
                  const line = `▸ Downloading ${p.percentage.toFixed(0)}% (${mb(p.downloaded)}/${mb(p.total)} MB)`;
                  process.stderr.write(process.stderr.isTTY ? `\r${line}` : `${line}\n`);
                  if (p.percentage >= 100)
                      process.stderr.write('\n');
              }
          });
          const samples = [
              'Machine learning is a subset of artificial intelligence that focuses on algorithms that can learn and make predictions from data without being explicitly programmed for every task.',
              'Deep learning uses neural networks with multiple layers to process and learn from complex data patterns, enabling breakthroughs in image recognition and natural language processing.',
              'Natural language processing combines computational linguistics with machine learning to help computers understand, interpret, and generate human language in a meaningful way.',
              'Computer vision enables machines to interpret and understand visual information from the world, using techniques like image classification, object detection, and facial recognition.',
              'Quantum computing leverages quantum mechanical phenomena to process information in fundamentally different ways than classical computers, potentially solving certain problems exponentially faster.',
              'Blockchain technology creates decentralized, immutable ledgers that enable secure peer-to-peer transactions without requiring a central authority or intermediary.',
              'Cloud computing delivers computing services over the internet, allowing users to access resources like storage, processing power, and applications on-demand from anywhere.',
              'Cybersecurity protects digital systems, networks, and data from malicious attacks, unauthorized access, and various forms of cyber threats through multiple layers of defense.'
          ];
          console.log('▸ Ingesting documents...');
          const result = await ragIngest({
              modelId,
              workspace,
              documents: samples,
              chunk: false
          });
          console.log(`▸ Ingested ${result.processed.length} documents`);
          console.log('▸ Searching for similar documents...');
          const results = await ragSearch({
              modelId,
              workspace,
              query,
              topK: 3
          });
          console.log('▸ Top 3 most similar documents:');
          results.forEach((result, index) => {
              console.log(`${index + 1}. (Score: ${result.score})`);
              console.log(`   ${result.content}`);
              console.log();
          });
          // Cleanup: close and delete workspace
          await ragCloseWorkspace({ workspace, deleteOnClose: true });
          console.log(`▸ Deleted '${workspace}' workspace`);
          await unloadModel({ modelId });
      }
      catch (error) {
          console.error('✖', error);
          process.exit(1);
      }
      ```
    </WrapCode>
  </Tab>

  <Tab value="ts" label="TypeScript">
    <WrapCode>
      ```ts file=<rootDir>/packages/sdk/examples/rag/rag-hyperdb/ingest.ts title="rag-managed.ts" lineNumbers
      import {
        loadModel,
        unloadModel,
        GTE_LARGE_FP16,
        ragIngest,
        ragSearch,
        ragCloseWorkspace
      } from '@qvac/sdk'

      try {
        // Get query from command line or use default
        const query = process.argv[2] || 'machine learning algorithms'
        const workspace = 'ingest-example'

        console.log(`▸ Query: "${query}"`)
        const modelId = await loadModel({
          modelSrc: GTE_LARGE_FP16,
          onProgress: (p) => {
            const mb = (n: number) => (n / 1e6).toFixed(1)
            const line = `▸ Downloading ${p.percentage.toFixed(0)}% (${mb(p.downloaded)}/${mb(p.total)} MB)`
            process.stderr.write(process.stderr.isTTY ? `\r${line}` : `${line}\n`)
            if (p.percentage >= 100) process.stderr.write('\n')
          }
        })

        const samples = [
          'Machine learning is a subset of artificial intelligence that focuses on algorithms that can learn and make predictions from data without being explicitly programmed for every task.',
          'Deep learning uses neural networks with multiple layers to process and learn from complex data patterns, enabling breakthroughs in image recognition and natural language processing.',
          'Natural language processing combines computational linguistics with machine learning to help computers understand, interpret, and generate human language in a meaningful way.',
          'Computer vision enables machines to interpret and understand visual information from the world, using techniques like image classification, object detection, and facial recognition.',
          'Quantum computing leverages quantum mechanical phenomena to process information in fundamentally different ways than classical computers, potentially solving certain problems exponentially faster.',
          'Blockchain technology creates decentralized, immutable ledgers that enable secure peer-to-peer transactions without requiring a central authority or intermediary.',
          'Cloud computing delivers computing services over the internet, allowing users to access resources like storage, processing power, and applications on-demand from anywhere.',
          'Cybersecurity protects digital systems, networks, and data from malicious attacks, unauthorized access, and various forms of cyber threats through multiple layers of defense.'
        ]

        console.log('▸ Ingesting documents...')
        const result = await ragIngest({
          modelId,
          workspace,
          documents: samples,
          chunk: false
        })
        console.log(`▸ Ingested ${result.processed.length} documents`)

        console.log('▸ Searching for similar documents...')
        const results = await ragSearch({
          modelId,
          workspace,
          query,
          topK: 3
        })

        console.log('▸ Top 3 most similar documents:')
        results.forEach((result, index) => {
          console.log(`${index + 1}. (Score: ${result.score})`)
          console.log(`   ${result.content}`)
          console.log()
        })

        // Cleanup: close and delete workspace
        await ragCloseWorkspace({ workspace, deleteOnClose: true })
        console.log(`▸ Deleted '${workspace}' workspace`)

        await unloadModel({ modelId })
      } catch (error) {
        console.error('✖', error)
        process.exit(1)
      }
      ```
    </WrapCode>
  </Tab>
</Tabs>

<Callout type="info">
  **Swap the backend:** to store the vectors in a TurboVec index instead of the default HyperDB, set `"ragTurbovec": true` in `qvac.config.json`. The RAG API is the same. See [Configuration › `ragTurbovec`](/sdk/v0.18/configuration#options) for details.
</Callout>

### TurboVec index with your own document store

The following script keeps the documents in a plain `Map`, embeds them with `embed()`, and indexes the vectors in a TurboVec index created with `createVectorIndex()`. Search returns the ids you assigned, so the documents can live in any database. The index is saved with `write()` and reopened with `loadVectorIndex()`; only the vectors are stored, never the documents.

<Tabs>
  <Tab value="js" label="JavaScript" default>
    <WrapCode>
      ```js file=<rootDir>/packages/sdk/dist/examples/rag/rag-turbovec.js title="rag-turbovec.js" lineNumbers
      import { createVectorIndex, embed, loadModel, loadVectorIndex, unloadModel, GTE_LARGE_FP16, VectorIndexStorage } from '@qvac/sdk';
      // Retrieval over documents kept in your own store. The SDK holds only the
      // vectors, in a TurboVec index inside the worker; the documents stay in this
      // Map (or any database you choose) and result ids map back to them.
      try {
          const query = process.argv[2] || 'Which moon has methane rain and lakes?';
          console.log(`▸ Query: "${query}"`);
          const documents = new Map([
              ['1', 'Saturn moon Titan has lakes, clouds, and rain made of liquid methane.'],
              ['2', 'Solar panels convert sunlight into electricity using photovoltaic cells.'],
              ['3', 'Honeybees communicate the location of flowers through a waggle dance.'],
              ['4', 'The Pacific Ocean is the largest and deepest ocean on Earth.']
          ]);
          const modelId = await loadModel({
              modelSrc: GTE_LARGE_FP16,
              onProgress: (p) => {
                  const mb = (n) => (n / 1e6).toFixed(1);
                  const line = `▸ Downloading ${p.percentage.toFixed(0)}% (${mb(p.downloaded)}/${mb(p.total)} MB)`;
                  process.stderr.write(process.stderr.isTTY ? `\r${line}` : `${line}\n`);
                  if (p.percentage >= 100)
                      process.stderr.write('\n');
              }
          });
          console.log('▸ Embedding documents...');
          const { embedding: vectors } = await embed({ modelId, text: [...documents.values()] });
          console.log('▸ Building the vector index...');
          const index = await createVectorIndex({
              dim: vectors[0].length,
              storage: VectorIndexStorage.TURBOVEC_Q4
          });
          await index.add({ ids: [...documents.keys()], vectors });
          console.log(`▸ Indexed ${index.length} vectors of dimension ${index.dim}`);
          console.log('▸ Searching...');
          const { embedding: queryVector } = await embed({ modelId, text: query });
          const hits = await index.search({ query: queryVector, k: 2 });
          console.log('▸ Top matches:');
          for (const hit of hits) {
              console.log(`  score=${hit.score.toFixed(4)} id=${hit.id}: ${documents.get(hit.id)}`);
          }
          // Snapshots persist the vectors only. A relative path resolves under the
          // QVAC data directory; pass an absolute path to store it elsewhere.
          const snapshotPath = 'examples/rag-turbovec.qvi';
          const { path: writtenPath } = await index.write({ path: snapshotPath });
          await index.dispose();
          console.log(`▸ Snapshot written to ${writtenPath}`);
          const reloaded = await loadVectorIndex({ path: snapshotPath });
          const [reloadedBest] = await reloaded.search({ query: queryVector, k: 1 });
          await reloaded.dispose();
          if (!reloadedBest || reloadedBest.id !== hits[0]?.id) {
              throw new Error(`Reloaded index returned id=${reloadedBest?.id} but the original returned id=${hits[0]?.id}`);
          }
          console.log(`▸ Reloaded index agrees: best match id=${reloadedBest.id}`);
          await unloadModel({ modelId });
      }
      catch (error) {
          console.error('✖', error);
          process.exit(1);
      }
      ```
    </WrapCode>
  </Tab>

  <Tab value="ts" label="TypeScript">
    <WrapCode>
      ```ts file=<rootDir>/packages/sdk/examples/rag/rag-turbovec.ts title="rag-turbovec.ts" lineNumbers
      import {
        createVectorIndex,
        embed,
        loadModel,
        loadVectorIndex,
        unloadModel,
        GTE_LARGE_FP16,
        VectorIndexStorage
      } from '@qvac/sdk'

      // Retrieval over documents kept in your own store. The SDK holds only the
      // vectors, in a TurboVec index inside the worker; the documents stay in this
      // Map (or any database you choose) and result ids map back to them.
      try {
        const query = process.argv[2] || 'Which moon has methane rain and lakes?'
        console.log(`▸ Query: "${query}"`)

        const documents = new Map<string, string>([
          ['1', 'Saturn moon Titan has lakes, clouds, and rain made of liquid methane.'],
          ['2', 'Solar panels convert sunlight into electricity using photovoltaic cells.'],
          ['3', 'Honeybees communicate the location of flowers through a waggle dance.'],
          ['4', 'The Pacific Ocean is the largest and deepest ocean on Earth.']
        ])

        const modelId = await loadModel({
          modelSrc: GTE_LARGE_FP16,
          onProgress: (p) => {
            const mb = (n: number) => (n / 1e6).toFixed(1)
            const line = `▸ Downloading ${p.percentage.toFixed(0)}% (${mb(p.downloaded)}/${mb(p.total)} MB)`
            process.stderr.write(process.stderr.isTTY ? `\r${line}` : `${line}\n`)
            if (p.percentage >= 100) process.stderr.write('\n')
          }
        })

        console.log('▸ Embedding documents...')
        const { embedding: vectors } = await embed({ modelId, text: [...documents.values()] })

        console.log('▸ Building the vector index...')
        const index = await createVectorIndex({
          dim: vectors[0]!.length,
          storage: VectorIndexStorage.TURBOVEC_Q4
        })
        await index.add({ ids: [...documents.keys()], vectors })
        console.log(`▸ Indexed ${index.length} vectors of dimension ${index.dim}`)

        console.log('▸ Searching...')
        const { embedding: queryVector } = await embed({ modelId, text: query })
        const hits = await index.search({ query: queryVector, k: 2 })
        console.log('▸ Top matches:')
        for (const hit of hits) {
          console.log(`  score=${hit.score.toFixed(4)} id=${hit.id}: ${documents.get(hit.id)}`)
        }

        // Snapshots persist the vectors only. A relative path resolves under the
        // QVAC data directory; pass an absolute path to store it elsewhere.
        const snapshotPath = 'examples/rag-turbovec.qvi'
        const { path: writtenPath } = await index.write({ path: snapshotPath })
        await index.dispose()
        console.log(`▸ Snapshot written to ${writtenPath}`)

        const reloaded = await loadVectorIndex({ path: snapshotPath })
        const [reloadedBest] = await reloaded.search({ query: queryVector, k: 1 })
        await reloaded.dispose()
        if (!reloadedBest || reloadedBest.id !== hits[0]?.id) {
          throw new Error(
            `Reloaded index returned id=${reloadedBest?.id} but the original returned id=${hits[0]?.id}`
          )
        }
        console.log(`▸ Reloaded index agrees: best match id=${reloadedBest.id}`)

        await unloadModel({ modelId })
      } catch (error) {
        console.error('✖', error)
        process.exit(1)
      }
      ```
    </WrapCode>
  </Tab>
</Tabs>

<Callout type="info">
  TurboVec quantises vectors (`VectorIndexStorage.TURBOVEC_Q4` by default, `TURBOVEC_Q2` for half the memory) and needs a dimension divisible by 8 and no greater than 1024. `GTE_LARGE_FP16` produces 1024-dimensional vectors and fits. Search uses dot products, so L2-normalise vectors when you need cosine similarity. The index lives in the SDK worker; after a worker restart, reopen it from the snapshot.
</Callout>

### External vector DB

#### MongoDB

The following script stores each document in MongoDB alongside its vector, builds a vector search index over it, and retrieves matches with a `$vectorSearch` aggregation that [pre-filters](https://www.mongodb.com/docs/vector-search/query/aggregation-stages/vector-search-stage/#pre-filter-search-data) by category:

<Tabs>
  <Tab value="js" label="JavaScript" default>
    <WrapCode>
      ```js file=<rootDir>/packages/sdk/dist/examples/rag/rag-mongodb.js title="rag-mongodb.js" lineNumbers
      import { embed, loadModel, unloadModel, GTE_LARGE_FP16 } from '@qvac/sdk';
      import { MongoClient } from 'mongodb';
      const INDEX_NAME = 'documents_vector_index';
      const MONGODB_SETUP_INSTRUCTIONS = `
      ▸ This example needs a MongoDB deployment with Atlas Vector Search.

      One way to get one is to run it in Docker:

         docker run -p 27017:27017 --name atlas-local mongodb/mongodb-atlas-local

      For more details, visit: https://www.mongodb.com/docs/atlas/cli/current/atlas-cli-deploy-docker/
      `;
      async function initializeMongoClient() {
          // Replace with your own deployment's connection string if it is not the Docker one
          // https://www.mongodb.com/docs/manual/reference/connection-string/
          const client = new MongoClient('mongodb://localhost:27017/?directConnection=true');
          try {
              await client.connect();
              await client.db('admin').command({ ping: 1 });
              console.log('▸ Connected to MongoDB server');
              return client;
          }
          catch {
              console.error('✖ Failed to connect to MongoDB server');
              console.error('▸ Please ensure the server is running on localhost:27017');
              console.error(MONGODB_SETUP_INSTRUCTIONS);
              process.exit(1);
          }
      }
      function wait(ms) {
          return new Promise((resolve) => setTimeout(resolve, ms));
      }
      try {
          // Get query and category from command line or use defaults
          const query = process.argv[2] || 'machine learning algorithms';
          const category = process.argv[3] || 'ai';
          console.log(`▸ Query: "${query}" (category: "${category}")`);
          const client = await initializeMongoClient();
          const collection = client.db('qvac').collection('documents');
          const modelId = await loadModel({
              modelSrc: GTE_LARGE_FP16,
              onProgress: (p) => {
                  const mb = (n) => (n / 1e6).toFixed(1);
                  const line = `▸ Downloading ${p.percentage.toFixed(0)}% (${mb(p.downloaded)}/${mb(p.total)} MB)`;
                  process.stderr.write(process.stderr.isTTY ? `\r${line}` : `${line}\n`);
                  if (p.percentage >= 100)
                      process.stderr.write('\n');
              }
          });
          // Sample corpus, each document tagged with a category to filter on
          const samples = [
              {
                  id: 1,
                  category: 'ai',
                  text: 'Machine learning is a subset of artificial intelligence that focuses on algorithms that can learn and make predictions from data without being explicitly programmed for every task.'
              },
              {
                  id: 2,
                  category: 'ai',
                  text: 'Deep learning uses neural networks with multiple layers to process and learn from complex data patterns, enabling breakthroughs in image recognition and natural language processing.'
              },
              {
                  id: 3,
                  category: 'ai',
                  text: 'Natural language processing combines computational linguistics with machine learning to help computers understand, interpret, and generate human language in a meaningful way.'
              },
              {
                  id: 4,
                  category: 'ai',
                  text: 'Computer vision enables machines to interpret and understand visual information from the world, using techniques like image classification, object detection, and facial recognition.'
              },
              {
                  id: 5,
                  category: 'computing',
                  text: 'Quantum computing leverages quantum mechanical phenomena to process information in fundamentally different ways than classical computers, potentially solving certain problems exponentially faster.'
              },
              {
                  id: 6,
                  category: 'security',
                  text: 'Blockchain technology creates decentralized, immutable ledgers that enable secure peer-to-peer transactions without requiring a central authority or intermediary.'
              },
              {
                  id: 7,
                  category: 'computing',
                  text: 'Cloud computing delivers computing services over the internet, allowing users to access resources like storage, processing power, and applications on-demand from anywhere.'
              },
              {
                  id: 8,
                  category: 'security',
                  text: 'Cybersecurity protects digital systems, networks, and data from malicious attacks, unauthorized access, and various forms of cyber threats through multiple layers of defense.'
              }
          ];
          // (Re)create the collection
          try {
              await collection.drop();
          }
          catch (e) {
              console.warn(`▸ Collection didn't exist, no need to drop: ${String(e)}`);
          }
          // Embed and store documents
          console.log('▸ Embedding documents...');
          const documents = [];
          for (const sample of samples) {
              const { embedding } = await embed({ modelId, text: sample.text });
              documents.push({
                  id: sample.id,
                  category: sample.category,
                  text: sample.text,
                  embedding
              });
          }
          await collection.insertMany(documents);
          // numDimensions is fixed at index creation and must match the model: GTE Large is 1024
          console.log('▸ Creating vector search index...');
          await collection.createSearchIndex({
              name: INDEX_NAME,
              type: 'vectorSearch',
              definition: {
                  fields: [
                      {
                          type: 'vector',
                          path: 'embedding',
                          numDimensions: 1024,
                          similarity: 'cosine'
                      },
                      // A field must be indexed as a filter to be usable in a $vectorSearch filter
                      {
                          type: 'filter',
                          path: 'category'
                      }
                  ]
              }
          });
          // Index builds are asynchronous; querying too early returns no matches
          for (let attempt = 0; attempt < 60; attempt++) {
              const [index] = (await collection.listSearchIndexes(INDEX_NAME).toArray());
              if (index?.queryable)
                  break;
              if (attempt === 59)
                  throw new Error(`Index ${INDEX_NAME} did not become queryable`);
              await wait(1000);
          }
          console.log('▸ Searching for similar documents...');
          const { embedding: queryEmbedding } = await embed({ modelId, text: query });
          const results = await collection
              .aggregate([
              {
                  $vectorSearch: {
                      index: INDEX_NAME,
                      path: 'embedding',
                      queryVector: queryEmbedding,
                      filter: { category: { $eq: category } },
                      numCandidates: 100,
                      limit: 3
                  }
              },
              {
                  $project: {
                      _id: 0,
                      id: 1,
                      category: 1,
                      text: 1,
                      score: { $meta: 'vectorSearchScore' }
                  }
              }
          ])
              .toArray();
          console.log('▸ Top 3 most similar documents:');
          results.forEach((result, index) => {
              console.log(`${index + 1}. (Score: ${result.score.toFixed(4)}, Category: ${result.category})`);
              console.log(`   ${result.text}`);
              console.log();
          });
          await unloadModel({ modelId });
          await client.close();
      }
      catch (error) {
          console.error('✖', error);
          process.exit(1);
      }
      ```
    </WrapCode>
  </Tab>

  <Tab value="ts" label="TypeScript">
    <WrapCode>
      ```ts file=<rootDir>/packages/sdk/examples/rag/rag-mongodb.ts title="rag-mongodb.ts" lineNumbers
      import { embed, loadModel, unloadModel, GTE_LARGE_FP16 } from '@qvac/sdk'
      import { MongoClient } from 'mongodb'

      const INDEX_NAME = 'documents_vector_index'

      const MONGODB_SETUP_INSTRUCTIONS = `
      ▸ This example needs a MongoDB deployment with Atlas Vector Search.

      One way to get one is to run it in Docker:

         docker run -p 27017:27017 --name atlas-local mongodb/mongodb-atlas-local

      For more details, visit: https://www.mongodb.com/docs/atlas/cli/current/atlas-cli-deploy-docker/
      `

      async function initializeMongoClient() {
        // Replace with your own deployment's connection string if it is not the Docker one
        // https://www.mongodb.com/docs/manual/reference/connection-string/
        const client = new MongoClient('mongodb://localhost:27017/?directConnection=true')

        try {
          await client.connect()
          await client.db('admin').command({ ping: 1 })
          console.log('▸ Connected to MongoDB server')
          return client
        } catch {
          console.error('✖ Failed to connect to MongoDB server')
          console.error('▸ Please ensure the server is running on localhost:27017')
          console.error(MONGODB_SETUP_INSTRUCTIONS)
          process.exit(1)
        }
      }

      function wait(ms: number) {
        return new Promise((resolve) => setTimeout(resolve, ms))
      }

      try {
        // Get query and category from command line or use defaults
        const query = process.argv[2] || 'machine learning algorithms'
        const category = process.argv[3] || 'ai'
        console.log(`▸ Query: "${query}" (category: "${category}")`)

        const client = await initializeMongoClient()
        const collection = client.db('qvac').collection('documents')

        const modelId = await loadModel({
          modelSrc: GTE_LARGE_FP16,
          onProgress: (p) => {
            const mb = (n: number) => (n / 1e6).toFixed(1)
            const line = `▸ Downloading ${p.percentage.toFixed(0)}% (${mb(p.downloaded)}/${mb(p.total)} MB)`
            process.stderr.write(process.stderr.isTTY ? `\r${line}` : `${line}\n`)
            if (p.percentage >= 100) process.stderr.write('\n')
          }
        })

        // Sample corpus, each document tagged with a category to filter on
        const samples = [
          {
            id: 1,
            category: 'ai',
            text: 'Machine learning is a subset of artificial intelligence that focuses on algorithms that can learn and make predictions from data without being explicitly programmed for every task.'
          },
          {
            id: 2,
            category: 'ai',
            text: 'Deep learning uses neural networks with multiple layers to process and learn from complex data patterns, enabling breakthroughs in image recognition and natural language processing.'
          },
          {
            id: 3,
            category: 'ai',
            text: 'Natural language processing combines computational linguistics with machine learning to help computers understand, interpret, and generate human language in a meaningful way.'
          },
          {
            id: 4,
            category: 'ai',
            text: 'Computer vision enables machines to interpret and understand visual information from the world, using techniques like image classification, object detection, and facial recognition.'
          },
          {
            id: 5,
            category: 'computing',
            text: 'Quantum computing leverages quantum mechanical phenomena to process information in fundamentally different ways than classical computers, potentially solving certain problems exponentially faster.'
          },
          {
            id: 6,
            category: 'security',
            text: 'Blockchain technology creates decentralized, immutable ledgers that enable secure peer-to-peer transactions without requiring a central authority or intermediary.'
          },
          {
            id: 7,
            category: 'computing',
            text: 'Cloud computing delivers computing services over the internet, allowing users to access resources like storage, processing power, and applications on-demand from anywhere.'
          },
          {
            id: 8,
            category: 'security',
            text: 'Cybersecurity protects digital systems, networks, and data from malicious attacks, unauthorized access, and various forms of cyber threats through multiple layers of defense.'
          }
        ]

        // (Re)create the collection
        try {
          await collection.drop()
        } catch (e) {
          console.warn(`▸ Collection didn't exist, no need to drop: ${String(e)}`)
        }

        // Embed and store documents
        console.log('▸ Embedding documents...')
        const documents = []
        for (const sample of samples) {
          const { embedding } = await embed({ modelId, text: sample.text })
          documents.push({
            id: sample.id,
            category: sample.category,
            text: sample.text,
            embedding
          })
        }

        await collection.insertMany(documents)

        // numDimensions is fixed at index creation and must match the model: GTE Large is 1024
        console.log('▸ Creating vector search index...')
        await collection.createSearchIndex({
          name: INDEX_NAME,
          type: 'vectorSearch',
          definition: {
            fields: [
              {
                type: 'vector',
                path: 'embedding',
                numDimensions: 1024,
                similarity: 'cosine'
              },
              // A field must be indexed as a filter to be usable in a $vectorSearch filter
              {
                type: 'filter',
                path: 'category'
              }
            ]
          }
        })

        // Index builds are asynchronous; querying too early returns no matches
        for (let attempt = 0; attempt < 60; attempt++) {
          const [index] = (await collection.listSearchIndexes(INDEX_NAME).toArray()) as {
            queryable?: boolean
          }[]
          if (index?.queryable) break
          if (attempt === 59) throw new Error(`Index ${INDEX_NAME} did not become queryable`)
          await wait(1000)
        }

        console.log('▸ Searching for similar documents...')
        const { embedding: queryEmbedding } = await embed({ modelId, text: query })

        const results = await collection
          .aggregate<{ id: number; category: string; text: string; score: number }>([
            {
              $vectorSearch: {
                index: INDEX_NAME,
                path: 'embedding',
                queryVector: queryEmbedding,
                filter: { category: { $eq: category } },
                numCandidates: 100,
                limit: 3
              }
            },
            {
              $project: {
                _id: 0,
                id: 1,
                category: 1,
                text: 1,
                score: { $meta: 'vectorSearchScore' }
              }
            }
          ])
          .toArray()

        console.log('▸ Top 3 most similar documents:')
        results.forEach((result, index) => {
          console.log(`${index + 1}. (Score: ${result.score.toFixed(4)}, Category: ${result.category})`)
          console.log(`   ${result.text}`)
          console.log()
        })

        await unloadModel({ modelId })
        await client.close()
      } catch (error) {
        console.error('✖', error)
        process.exit(1)
      }
      ```
    </WrapCode>
  </Tab>
</Tabs>

#### SQLite

The following script shows the same workflow using `embed()` plus a SQLite vector index:

<Tabs>
  <Tab value="js" label="JavaScript" default>
    <WrapCode>
      ```js file=<rootDir>/packages/sdk/dist/examples/rag/rag-sqlite.js title="rag-sqlite.js" lineNumbers
      import { embed, loadModel, unloadModel, GTE_LARGE_FP16 } from '@qvac/sdk';
      import sqlite3InitModule from '@sqliteai/sqlite-wasm';
      try {
          // Get query from command line or use default
          const query = process.argv[2] || 'machine learning algorithms';
          console.log(`▸ Query: "${query}"`);
          // Initialize SQLite with vector extension
          const sqlite3 = await sqlite3InitModule();
          const db = new sqlite3.oo1.DB(':memory:', 'c');
          const modelId = await loadModel({
              modelSrc: GTE_LARGE_FP16,
              onProgress: (p) => {
                  const mb = (n) => (n / 1e6).toFixed(1);
                  const line = `▸ Downloading ${p.percentage.toFixed(0)}% (${mb(p.downloaded)}/${mb(p.total)} MB)`;
                  process.stderr.write(process.stderr.isTTY ? `\r${line}` : `${line}\n`);
                  if (p.percentage >= 100)
                      process.stderr.write('\n');
              }
          });
          const samples = [
              {
                  id: 1,
                  text: 'Machine learning is a subset of artificial intelligence that focuses on algorithms that can learn and make predictions from data without being explicitly programmed for every task.'
              },
              {
                  id: 2,
                  text: 'Deep learning uses neural networks with multiple layers to process and learn from complex data patterns, enabling breakthroughs in image recognition and natural language processing.'
              },
              {
                  id: 3,
                  text: 'Natural language processing combines computational linguistics with machine learning to help computers understand, interpret, and generate human language in a meaningful way.'
              },
              {
                  id: 4,
                  text: 'Computer vision enables machines to interpret and understand visual information from the world, using techniques like image classification, object detection, and facial recognition.'
              },
              {
                  id: 5,
                  text: 'Quantum computing leverages quantum mechanical phenomena to process information in fundamentally different ways than classical computers, potentially solving certain problems exponentially faster.'
              },
              {
                  id: 6,
                  text: 'Blockchain technology creates decentralized, immutable ledgers that enable secure peer-to-peer transactions without requiring a central authority or intermediary.'
              },
              {
                  id: 7,
                  text: 'Cloud computing delivers computing services over the internet, allowing users to access resources like storage, processing power, and applications on-demand from anywhere.'
              },
              {
                  id: 8,
                  text: 'Cybersecurity protects digital systems, networks, and data from malicious attacks, unauthorized access, and various forms of cyber threats through multiple layers of defense.'
              }
          ];
          // Create table for documents with vector storage
          db.exec(`
        CREATE TABLE IF NOT EXISTS documents (
          id INTEGER PRIMARY KEY,
          text TEXT NOT NULL,
          embedding BLOB NOT NULL
        )
      `);
          console.log('▸ Embedding documents...');
          for (const sample of samples) {
              const { embedding } = await embed({ modelId, text: sample.text });
              db.exec({
                  sql: 'INSERT INTO documents VALUES (?, ?, vector_as_f32(?))',
                  bind: [sample.id, sample.text, JSON.stringify(embedding)]
              });
          }
          // Initialize and optimize vector index
          db.exec(`SELECT vector_init('documents', 'embedding', 'type=FLOAT32,dimension=1024')`);
          // Quantize vectors
          db.exec(`SELECT vector_quantize('documents', 'embedding')`);
          // [Optional] Preload quantized vectors in memory for optimal performance
          db.exec(`SELECT vector_quantize_preload('documents', 'embedding')`);
          // Search for similar documents
          console.log('▸ Searching for similar documents...');
          const { embedding: queryEmbedding } = await embed({ modelId, text: query });
          const results = [];
          // Perform vector search
          db.exec({
              sql: `
          SELECT d.id, d.text, v.distance 
          FROM documents d
          JOIN vector_quantize_scan('documents', 'embedding', vector_as_f32(?), 3) v
          ON d.id = v.rowid
        `,
              bind: [JSON.stringify(queryEmbedding)],
              rowMode: 'object',
              callback: (row) => {
                  const typedRow = row;
                  results.push(typedRow);
              }
          });
          console.log('\n▸ Top 3 most similar documents:');
          results.forEach((result, index) => {
              console.log('='.repeat(50) + ' Top result:');
              console.log(`\n${index + 1}. [ID: ${result.id}] (Score: ${result.distance.toFixed(4)})`);
              console.log(`   ${result.text}`);
              console.log('='.repeat(100));
              console.log();
          });
          await unloadModel({ modelId });
          db.close();
      }
      catch (error) {
          console.error('✖', error);
          process.exit(1);
      }
      ```
    </WrapCode>
  </Tab>

  <Tab value="ts" label="TypeScript">
    <WrapCode>
      ```ts file=<rootDir>/packages/sdk/examples/rag/rag-sqlite.ts title="rag-sqlite.ts" lineNumbers
      import { embed, loadModel, unloadModel, GTE_LARGE_FP16 } from '@qvac/sdk'
      import sqlite3InitModule from '@sqliteai/sqlite-wasm'

      try {
        // Get query from command line or use default
        const query = process.argv[2] || 'machine learning algorithms'
        console.log(`▸ Query: "${query}"`)

        // Initialize SQLite with vector extension
        const sqlite3 = await sqlite3InitModule()
        const db = new sqlite3.oo1.DB(':memory:', 'c')

        const modelId = await loadModel({
          modelSrc: GTE_LARGE_FP16,
          onProgress: (p) => {
            const mb = (n: number) => (n / 1e6).toFixed(1)
            const line = `▸ Downloading ${p.percentage.toFixed(0)}% (${mb(p.downloaded)}/${mb(p.total)} MB)`
            process.stderr.write(process.stderr.isTTY ? `\r${line}` : `${line}\n`)
            if (p.percentage >= 100) process.stderr.write('\n')
          }
        })

        const samples = [
          {
            id: 1,
            text: 'Machine learning is a subset of artificial intelligence that focuses on algorithms that can learn and make predictions from data without being explicitly programmed for every task.'
          },
          {
            id: 2,
            text: 'Deep learning uses neural networks with multiple layers to process and learn from complex data patterns, enabling breakthroughs in image recognition and natural language processing.'
          },
          {
            id: 3,
            text: 'Natural language processing combines computational linguistics with machine learning to help computers understand, interpret, and generate human language in a meaningful way.'
          },
          {
            id: 4,
            text: 'Computer vision enables machines to interpret and understand visual information from the world, using techniques like image classification, object detection, and facial recognition.'
          },
          {
            id: 5,
            text: 'Quantum computing leverages quantum mechanical phenomena to process information in fundamentally different ways than classical computers, potentially solving certain problems exponentially faster.'
          },
          {
            id: 6,
            text: 'Blockchain technology creates decentralized, immutable ledgers that enable secure peer-to-peer transactions without requiring a central authority or intermediary.'
          },
          {
            id: 7,
            text: 'Cloud computing delivers computing services over the internet, allowing users to access resources like storage, processing power, and applications on-demand from anywhere.'
          },
          {
            id: 8,
            text: 'Cybersecurity protects digital systems, networks, and data from malicious attacks, unauthorized access, and various forms of cyber threats through multiple layers of defense.'
          }
        ]

        // Create table for documents with vector storage
        db.exec(`
        CREATE TABLE IF NOT EXISTS documents (
          id INTEGER PRIMARY KEY,
          text TEXT NOT NULL,
          embedding BLOB NOT NULL
        )
      `)

        console.log('▸ Embedding documents...')
        for (const sample of samples) {
          const { embedding } = await embed({ modelId, text: sample.text })
          db.exec({
            sql: 'INSERT INTO documents VALUES (?, ?, vector_as_f32(?))',
            bind: [sample.id, sample.text, JSON.stringify(embedding)]
          })
        }

        // Initialize and optimize vector index
        db.exec(`SELECT vector_init('documents', 'embedding', 'type=FLOAT32,dimension=1024')`)

        // Quantize vectors
        db.exec(`SELECT vector_quantize('documents', 'embedding')`)

        // [Optional] Preload quantized vectors in memory for optimal performance
        db.exec(`SELECT vector_quantize_preload('documents', 'embedding')`)

        // Search for similar documents
        console.log('▸ Searching for similar documents...')
        const { embedding: queryEmbedding } = await embed({ modelId, text: query })

        const results: Array<{
          id: number
          text: string
          distance: number
        }> = []

        // Perform vector search
        db.exec({
          sql: `
          SELECT d.id, d.text, v.distance 
          FROM documents d
          JOIN vector_quantize_scan('documents', 'embedding', vector_as_f32(?), 3) v
          ON d.id = v.rowid
        `,
          bind: [JSON.stringify(queryEmbedding)],
          rowMode: 'object',
          callback: (row: unknown) => {
            const typedRow = row as { id: number; text: string; distance: number }
            results.push(typedRow)
          }
        })

        console.log('\n▸ Top 3 most similar documents:')
        results.forEach((result, index) => {
          console.log('='.repeat(50) + ' Top result:')
          console.log(`\n${index + 1}. [ID: ${result.id}] (Score: ${result.distance.toFixed(4)})`)
          console.log(`   ${result.text}`)
          console.log('='.repeat(100))
          console.log()
        })

        await unloadModel({ modelId })
        db.close()
      } catch (error) {
        console.error('✖', error)
        process.exit(1)
      }
      ```
    </WrapCode>
  </Tab>
</Tabs>

<Callout type="info">
  The Python client supports this capability through the same worker. A dedicated Python example is not yet published — see the [Python SDK](/sdk/v0.18/python-sdk) for the API surface.
</Callout>

<Callout type="success">
  **Tip:** all examples throughout this documentation are self-contained and runnable. For instructions on how to run them, see the [JS/TS quickstart](/sdk/v0.18/js-ts-sdk#quickstart) or the [Python quickstart](/sdk/v0.18/python-sdk#quickstart).
</Callout>
