Illustration of a brain-shaped circuit inside a locked device representing local LLM inference
Run LLM inference locally on edge devices with WebAssembly and ONNX Runtime

Building Local-First RAG Systems: Running Privacy-Preserving LLMs on Edge Devices with WebAssembly and ONNX Runtime

Practical guide to building local-first RAG systems with WebAssembly and ONNX Runtime—run privacy-preserving LLMs on edge devices with performance tips and code.

Building Local-First RAG Systems: Running Privacy-Preserving LLMs on Edge Devices with WebAssembly and ONNX Runtime

Introduction

Developers are building systems that must answer queries over private data without sending sensitive information to cloud APIs. Retrieval-Augmented Generation (RAG) is ideal for these use cases, but most RAG deployments rely on remote LLMs. This post shows a concrete, practical approach to building local-first RAG systems that run inference on-device using WebAssembly (Wasm) and ONNX Runtime. You’ll get an architecture overview, model-selection guidance, a runnable code pattern for embedding + retrieval + generation, and performance tuning tips for edge environments.

Why local-first RAG?

Trade-offs to accept upfront:

Architecture: components and flow

High-level components for a local-first RAG system:

  1. Document ingestion: chunk and store content locally (SQLite, local file, or IndexedDB).
  2. Embedding model: produce embeddings for documents and queries using an on-device model exported to ONNX.
  3. Vector store: light-weight nearest-neighbor search (HNSW, brute-force for small corpora) running in Wasm or native.
  4. Retriever: fetch top-k passages for a query.
  5. Generator LLM: an on-device model that conditions on retrieved context and generates answers.
  6. Orchestration: glue code that runs embeddings, retrieval, and generation, respecting memory and CPU limits.

The simplest flow: ingest → embed → index → query-embed → retrieve → generate.

Choosing models and quantization

Practical selection rules:

Be explicit about thresholds: models 3B (3B) often require optimizations; models 7B likely won’t fit on constrained devices without offloading.

Why ONNX Runtime + WebAssembly?

Small note: where native inference is available (mobile NN APIs, local server), consider ORT native backends. But Wasm unlocks the broadest portability.

Practical integration: embedding + retrieval + generation

You’ll typically run two ONNX sessions: one for the embedder and one for the generator. Keep them separate to reduce peak memory. The code pattern below demonstrates the minimal flow in the browser/Electron using onnxruntime-web.

4-space indented blocks are used for multi-line code.

import * as ort from 'onnxruntime-web'

// Initialize Wasm execution provider
await ort.InferenceSession.create('/models/embedder.onnx', 
    { executionProviders: ['wasm'] })

async function embedText(session, inputTokens) {
    // Prepare input tensor depending on model requirements
    const feeds = { input_ids: new ort.Tensor('int32', inputTokens, [1, inputTokens.length]) }
    const results = await session.run(feeds)
    return results.last_hidden_state.data
}

// Simple cosine similarity for ranking
function cosine(a, b) {
    let dot = 0.0, na = 0.0, nb = 0.0
    for (let i = 0; i < a.length; i++) {
        dot += a[i] * b[i]
        na += a[i] * a[i]
        nb += b[i] * b[i]
    }
    return dot / (Math.sqrt(na) * Math.sqrt(nb) + 1e-10)
}

This example is intentionally minimal: real embedder inputs must match the tokenization and shape your ONNX model expects.

Indexing and retrieval choices

Example pseudo-flow for retrieval:

// embeddings: Array of Float32Array
// queryEmbedding: Float32Array
function topK(embeddings, queryEmbedding, k) {
    const scores = embeddings.map(e => cosine(e, queryEmbedding))
    const idx = scores
        .map((s, i) => ({ s, i }))
        .sort((a, b) => b.s - a.s)
        .slice(0, k)
        .map(x => x.i)
    return idx
}

Generation: conditioning on retrieved context

To keep models small, you can use a prompt template that concatenates retrieved passages and instructs the model to answer concisely. Watch prompt length; if the total token count of retrieved context plus the prompt exceeds model capacity, truncate documents by relevance or use passage reranking.

Key implementation detail: streaming decoding is critical for perceived latency. ONNX Runtime supports operator-level execution; implement a token-by-token loop that feeds the previous tokens into the model and appends outputs, or use a precompiled sampling kernel where available.

Performance tuning and memory budgeting

Be explicit about constraints: set a hard memory budget for the runtime and measure peak RSS on target hardware.

Security and privacy considerations

Deployment patterns

When updating models, use a staged rollout to validate model behavior on-device before broad distribution.

Example checklist (summary)

Final notes

Local-first RAG systems trade the convenience of remote LLM APIs for privacy and predictable latency. WebAssembly and ONNX Runtime make this practical today across browsers, desktops, and embedded devices. The engineering work focuses on model selection, efficient indexing, and careful memory budgeting. Start small: embed a subset of data, test retrieval and local generation, and iterate on quantization and index strategies.

Build incrementally, measure on-device, and keep security at the center: when done right, local-first RAG delivers strong privacy guarantees and responsive UX without cloud dependencies.

Related

Get sharp weekly insights

Newsletter coming soon. Stay tuned for curated deep dives on edge AI and autonomous systems.