---
title: "@qvac/llm-llamacpp"
canonical: https://docs.qvac.tether.io/ecosystem/addons/llm-llamacpp/
collection: "Ecosystem"
---

# @qvac/llm-llamacpp (/ecosystem/addons/llm-llamacpp)



## Overview

[Bare module](https://bare.pears.com) that adds support for text completion and multimodal prompts in QVAC using [`qvac-fabric-llm.cpp`](https://github.com/tetherto/qvac-fabric-llm.cpp) as the inference engine.

## Models

You can load any [`llama.cpp`](https://github.com/ggml-org/llama.cpp)-compatible text-generation/chat model. Model file format: `*.gguf`.

## Requirement

Bare `>= v1.24`

## Installation

```bash
npm i @qvac/llm-llamacpp
```

`@qvac/llm-llamacpp` links the shared `@qvac/fabric` runtime. `@qvac/fabric`
ships its native runtime and ggml backends for each desktop host through a
version-locked, `os`/`cpu` filtered optional dependency:

| Host                | Package                     |
| ------------------- | --------------------------- |
| linux-x64 (glibc)   | `@qvac/fabric-linux-x64`    |
| linux-arm64 (glibc) | `@qvac/fabric-linux-arm64`  |
| darwin-arm64        | `@qvac/fabric-darwin-arm64` |
| darwin-x64          | `@qvac/fabric-darwin-x64`   |
| win32-x64           | `@qvac/fabric-win32-x64`    |

Do not depend on the desktop platform packages directly. Supported installers
are npm 7+, pnpm, bun, and Yarn Berry; Yarn v1 and `--omit=optional` installs
skip the platform package and fail at require time with an error naming it.

Mobile targets are cross-built, so no install host ever matches their `os`,
and optional-dependency filtering can never select them. Mobile applications
must declare `@qvac/fabric` and the target's fabric platform package as direct
dependencies at the same exact version, one that satisfies the `@qvac/fabric`
range `@qvac/llm-llamacpp` declares:

| Target                    | Package                      |
| ------------------------- | ---------------------------- |
| android-arm64             | `@qvac/fabric-android-arm64` |
| ios (device + simulators) | `@qvac/fabric-ios`           |

```json
{
  "dependencies": {
    "@qvac/llm-llamacpp": "x.y.z",
    "@qvac/fabric": "a.b.c",
    "@qvac/fabric-android-arm64": "a.b.c"
  }
}
```

## Quickstart

<Steps>
  <Step>
    If you don't have Bare runtime, install it:

    ```bash
    npm i -g bare
    ```
  </Step>

  <Step>
    Create a new project:

    ```bash
    mkdir qvac-llm-quickstart
    cd qvac-llm-quickstart
    npm init -y
    ```
  </Step>

  <Step>
    Install dependencies:

    ```bash
    npm i @qvac/llm-llamacpp bare-path bare-process
    ```
  </Step>

  <Step>
    Download a compatible model:

    ```bash
    curl -L --create-dirs -o models/Llama-3.2-1B-Instruct-Q4_0.gguf \
      https://huggingface.co/bartowski/Llama-3.2-1B-Instruct-GGUF/resolve/main/Llama-3.2-1B-Instruct-Q4_0.gguf
    ```
  </Step>

  <Step>
    Create `index.js`:
  </Step>

  <WrapCode>
    ```js title="index.js" lineNumbers
    'use strict'

    const LlmLlamacpp = require('@qvac/llm-llamacpp')
    const path = require('bare-path')
    const process = require('bare-process')

    async function main () {
      const modelName = 'Llama-3.2-1B-Instruct-Q4_0.gguf'
      const dirPath = path.resolve('./models')
      const modelPath = path.join(dirPath, modelName)

      // 1. Configuring model settings
      const config = {
        device: 'gpu',
        gpu_layers: '999',
        ctx_size: '1024'
      }

      // 2. Loading model
      const model = new LlmLlamacpp({
        files: { model: [modelPath] },
        config,
        logger: console,
        opts: { stats: true }
      })
      await model.load()

      try {
        // 3. Running inference with conversation prompt
        const prompt = [
          {
            role: 'system',
            content: 'You are a helpful, respectful and honest assistant.'
          },
          {
            role: 'user',
            content: 'what is bitcoin?'
          },
          {
            role: 'assistant',
            content: "It's a digital currency."
          },
          {
            role: 'user',
            content: 'Can you elaborate on the previous topic?'
          }
        ]

        const response = await model.run(prompt)
        let fullResponse = ''

        await response
          .onUpdate(data => {
            process.stdout.write(data)
            fullResponse += data
          })
          .await()

        console.log('\n')
        console.log('Full response:\n', fullResponse)
        console.log(`Inference stats: ${JSON.stringify(response.stats)}`)
      } catch (error) {
        const errorMessage = error?.message || error?.toString() || String(error)
        console.error('Error occurred:', errorMessage)
        console.error('Error details:', error)
      } finally {
        // 4. Cleaning up resources
        await model.unload()
      }
    }

    main().catch(error => {
      console.error('Fatal error in main function:', {
        error: error.message,
        stack: error.stack,
        timestamp: new Date().toISOString()
      })
      process.exit(1)
    })
    ```
  </WrapCode>

  <Step>
    Run `index.js`:

    ```bash
    bare index.js
    ```
  </Step>
</Steps>

## Usage

### 1. Import the Model Class

```js
const LlmLlamacpp = require('@qvac/llm-llamacpp')
```

### 2. Create Local Model Paths

The addon reads GGUF files directly from disk. Download the model, then pass absolute local paths to `files.model`.

```js
const path = require('bare-path')

const dirPath = path.resolve('./models')
const modelName = 'Llama-3.2-1B-Instruct-Q4_0.gguf'

const modelPath = path.join(dirPath, modelName)
```

### 3. Create the `args` obj

```js
// a minimal config; see step 4 for all available options
const config = {
  gpu_layers: '99',
  ctx_size: '1024',
  device: 'cpu'
}

const args = {
  files: {
    model: [modelPath],
    // projectionModel: path.join(dirPath, 'mmproj-SmolVLM2-500M-Video-Instruct-Q8_0.gguf') // for multimodal support pass the projection model path
  },
  config,
  opts: { stats: true },
  logger: console
}
```

The `args` obj contains the following properties:

* `files.model`: Required. An array of absolute paths to the GGUF model file(s) to load. For sharded models, provide every shard and companion file in order.
* `files.projectionModel`: Optional. Absolute path to the projection model file. This is required for multimodal support.
* `config`: The model configuration object.
* `logger`: This property is used to create a `QvacLogger` instance, which handles all logging functionality.
* `opts.stats`: This flag determines whether to calculate inference stats.

### 4. Create the `config` obj

The `config` obj consists of a set of hyper-parameters which can be used to tweak the behaviour of the model.\
*All parameters must be strings.*

```js
// an example of possible configuration
const config = {
  gpu_layers: '99', // number of model layers offloaded to GPU.
  ctx_size: '1024', // context length
  device: 'cpu', // must be specified: 'gpu' or 'cpu' else it will throw an error
  load_mode: 'none' // read fully into memory: no mmap, mlock or direct I/O
}
```

| Parameter          | Range / Type                                              | Default                      | Description                                        |
| ------------------ | --------------------------------------------------------- | ---------------------------- | -------------------------------------------------- |
| device             | `"gpu"` or `"cpu"`                                        | — (required)                 | Device to run inference on                         |
| gpu\_layers        | integer                                                   | 0                            | Number of model layers to offload to GPU           |
| ctx\_size          | 0 – model-dependent                                       | 4096 (0 = loaded from model) | Context window size                                |
| lora               | string                                                    | —                            | Path to LoRA adapter file                          |
| temp               | 0.00 – 2.00                                               | 0.8                          | Sampling temperature                               |
| top\_p             | 0 – 1                                                     | 0.9                          | Top-p (nucleus) sampling                           |
| top\_k             | 0 – 128                                                   | 40                           | Top-k sampling                                     |
| predict            | integer (-1 = infinity)                                   | -1                           | Maximum tokens to predict                          |
| seed               | integer                                                   | -1 (random)                  | Random seed for sampling                           |
| load\_mode         | `"none"`, `"mmap"`, `"mlock"`, `"mmap+mlock"`, or `"dio"` | `"mmap"`                     | Select the model loading mode                      |
| reverse\_prompt    | string (comma-separated)                                  | —                            | Stop generation when these strings are encountered |
| repeat\_penalty    | float                                                     | 1.1                          | Repetition penalty                                 |
| presence\_penalty  | float                                                     | 0                            | Presence penalty for sampling                      |
| frequency\_penalty | float                                                     | 0                            | Frequency penalty for sampling                     |
| tools              | `"true"` or `"false"`                                     | `"false"`                    | Enable tool calling with jinja templating          |
| verbosity          | 0 – 3 (0=ERROR, 1=WARNING, 2=INFO, 3=DEBUG)               | 0                            | Logging verbosity level                            |
| main-gpu           | integer, `"integrated"`, or `"dedicated"`                 | —                            | GPU selection for multi-GPU systems                |

#### IGPU/GPU  selection logic:

| Scenario                       | main-gpu not specified            | main-gpu: `"dedicated"` | main-gpu: `"integrated"` |
| ------------------------------ | --------------------------------- | ----------------------- | ------------------------ |
| Devices considered             | All GPUs (dedicated + integrated) | Only dedicated GPUs     | Only integrated GPUs     |
| System with iGPU only          | ✅ Uses iGPU                       | ❌ Falls back to CPU     | ✅ Uses iGPU              |
| System with dedicated GPU only | ✅ Uses dedicated GPU              | ✅ Uses dedicated GPU    | ❌ Falls back to CPU      |
| System with both               | ✅ Uses dedicated GPU (preferred)  | ✅ Uses dedicated GPU    | ✅ Uses integrated GPU    |

### 5. Create Model Instance

```js
const model = new LlmLlamacpp(args)
```

### 6. Load Model

```js
await model.load()
```

Loads the model file(s) passed in `files.model` and activates the native addon. If a projection model was provided (`files.projectionModel`), it is loaded as part of the same step.

### 7. Run Inference

Pass an array of messages (following the chat completion format) to the `run` method. Process the generated tokens asynchronously:

```javascript
try {
  const messages = [
    { role: 'system', content: 'You are a helpful assistant.' },
    { role: 'user', content: 'What is the capital of France?' }
  ]

  const response = await model.run(messages)
  const buffer = []

  // Option 1: Process streamed output using async iterator
  for await (const token of response.iterate()) {
    process.stdout.write(token) // Write token directly to output
    buffer.push(token)
  }

  // Option 2: Process streamed output using callback
  await response.onUpdate(token => ).await()

  console.log('\n--- Full Response ---\n', buffer.join(''))

} catch (error) {
  console.error('Inference failed:', error)
}
```

### 8. Release Resources

Unload the model when finished:

```javascript
try {
  await model.unload()
} catch (error) {
  console.error('Failed to unload model:', error)
}
```

## More resources

[Package at npm](https://www.npmjs.com/package/@qvac/llm-llamacpp)
