Building Local-First RAG Systems: Running Privacy-Preserving LLMs on Edge Devices with WebAssembly and ONNX Runtime
Practical guide to building local-first RAG systems with WebAssembly and ONNX Runtime—run privacy-preserving LLMs on edge devices with performance tips and code.
Building Local-First RAG Systems: Running Privacy-Preserving LLMs on Edge Devices with WebAssembly and ONNX Runtime
Introduction
Developers are building systems that must answer queries over private data without sending sensitive information to cloud APIs. Retrieval-Augmented Generation (RAG) is ideal for these use cases, but most RAG deployments rely on remote LLMs. This post shows a concrete, practical approach to building local-first RAG systems that run inference on-device using WebAssembly (Wasm) and ONNX Runtime. You’ll get an architecture overview, model-selection guidance, a runnable code pattern for embedding + retrieval + generation, and performance tuning tips for edge environments.
Why local-first RAG?
- Privacy: Data never leaves the device. No RPCs to third-party LLMs.
- Predictable latency: Inference latency becomes local and deterministic.
- Offline capability: Crucial for field devices, air-gapped systems, and regulated industries.
Trade-offs to accept upfront:
- Model size and compute limits: Edge devices often have constrained RAM/CPU. Plan for quantized models or smaller architectures.
- Maintenance: Updating models requires shipping new binaries or in-app updates.
Architecture: components and flow
High-level components for a local-first RAG system:
- Document ingestion: chunk and store content locally (SQLite, local file, or IndexedDB).
- Embedding model: produce embeddings for documents and queries using an on-device model exported to ONNX.
- Vector store: light-weight nearest-neighbor search (HNSW, brute-force for small corpora) running in Wasm or native.
- Retriever: fetch top-k passages for a query.
- Generator LLM: an on-device model that conditions on retrieved context and generates answers.
- Orchestration: glue code that runs embeddings, retrieval, and generation, respecting memory and CPU limits.
The simplest flow: ingest → embed → index → query-embed → retrieve → generate.
Choosing models and quantization
Practical selection rules:
- Use small generator models if you need sub-second responses on mobile. Distillations and 3B-7B models are common trade-offs. For strict edge constraints, prefer 1.4B-class or purpose-built smaller decoders.
- Embedder size: embedding models can be smaller (sentence-transformers 384/768 dims). Lower dims reduce memory and index size.
- Quantize aggressively: 8-bit, 4-bit, or mixed precision reduce memory and speed up Wasm inference. Convert PyTorch/TensorFlow to ONNX then apply quantization.
Be explicit about thresholds: models 3B (3B) often require optimizations; models 7B likely won’t fit on constrained devices without offloading.
Why ONNX Runtime + WebAssembly?
- ONNX Runtime (ORT) supports exporting many model types to ONNX and running them in WebAssembly using
onnxruntime-web. - Wasm runs in browsers, Electron, and many embedded runtimes. It provides portable dependency-free execution.
- ORT Wasm can use multi-threaded Wasm supported by workers and WASI where available.
Small note: where native inference is available (mobile NN APIs, local server), consider ORT native backends. But Wasm unlocks the broadest portability.
Practical integration: embedding + retrieval + generation
You’ll typically run two ONNX sessions: one for the embedder and one for the generator. Keep them separate to reduce peak memory. The code pattern below demonstrates the minimal flow in the browser/Electron using onnxruntime-web.
4-space indented blocks are used for multi-line code.
import * as ort from 'onnxruntime-web'
// Initialize Wasm execution provider
await ort.InferenceSession.create('/models/embedder.onnx',
{ executionProviders: ['wasm'] })
async function embedText(session, inputTokens) {
// Prepare input tensor depending on model requirements
const feeds = { input_ids: new ort.Tensor('int32', inputTokens, [1, inputTokens.length]) }
const results = await session.run(feeds)
return results.last_hidden_state.data
}
// Simple cosine similarity for ranking
function cosine(a, b) {
let dot = 0.0, na = 0.0, nb = 0.0
for (let i = 0; i < a.length; i++) {
dot += a[i] * b[i]
na += a[i] * a[i]
nb += b[i] * b[i]
}
return dot / (Math.sqrt(na) * Math.sqrt(nb) + 1e-10)
}
This example is intentionally minimal: real embedder inputs must match the tokenization and shape your ONNX model expects.
Indexing and retrieval choices
- For small corpora (<10k documents), CPU brute-force search with optimized SIMD vectors in Wasm is acceptable.
- For larger corpora, compile HNSW or a compact ANN engine to Wasm (projects like
hnswlibhave Wasm ports) and persist the index to IndexedDB or local filesystem.
Example pseudo-flow for retrieval:
// embeddings: Array of Float32Array
// queryEmbedding: Float32Array
function topK(embeddings, queryEmbedding, k) {
const scores = embeddings.map(e => cosine(e, queryEmbedding))
const idx = scores
.map((s, i) => ({ s, i }))
.sort((a, b) => b.s - a.s)
.slice(0, k)
.map(x => x.i)
return idx
}
Generation: conditioning on retrieved context
To keep models small, you can use a prompt template that concatenates retrieved passages and instructs the model to answer concisely. Watch prompt length; if the total token count of retrieved context plus the prompt exceeds model capacity, truncate documents by relevance or use passage reranking.
Key implementation detail: streaming decoding is critical for perceived latency. ONNX Runtime supports operator-level execution; implement a token-by-token loop that feeds the previous tokens into the model and appends outputs, or use a precompiled sampling kernel where available.
Performance tuning and memory budgeting
- Memory: load only one session at a time if memory is tight. Load embedder, build index, then unload and load generator only when needed.
- Threads: enable Wasm threads where supported to exploit multi-core.
- Quantization: test 8-bit vs 4-bit—sometimes 4-bit reduces quality; verify with real prompts.
- Batch embeddings for ingestion; use streaming for inference.
Be explicit about constraints: set a hard memory budget for the runtime and measure peak RSS on target hardware.
Security and privacy considerations
- All inference artifacts (models, indices) stored on-device must be encrypted if the device is shared.
- Protect model weights anti-tamper if needed; adversaries could exfiltrate sensitive training data from model parameters in extreme cases.
- Logging: avoid writing raw user queries to disk. If you log, anonymize or store only non-sensitive metadata.
Deployment patterns
- Browser-based apps: bundle Wasm + ONNX models as static assets and load with
onnxruntime-web. Persist indices in IndexedDB. - Desktop (Electron): same as browser but you can use the native ONNX Runtime for better perf when available.
- Embedded Linux: use Wasm runtimes (Wasmtime) or native ORT builds for better performance.
When updating models, use a staged rollout to validate model behavior on-device before broad distribution.
Example checklist (summary)
- Choose models sized for your device: smaller generator, compact embedder.
- Convert and quantize models to ONNX; test locally for quality.
- Use
onnxruntime-webin Wasm for portability; enable threads if available. - Store document embeddings and indices locally (IndexedDB, SQLite, or filesystem).
- Implement retrieval (HNSW/Wasm or brute-force) and reranking to control prompt size.
- Stream decoding for better UX; limit context length and truncate by relevance.
- Encrypt on-disk artifacts and minimize logging of raw queries.
- Test latency and memory on target hardware; tune quantization and threading.
Final notes
Local-first RAG systems trade the convenience of remote LLM APIs for privacy and predictable latency. WebAssembly and ONNX Runtime make this practical today across browsers, desktops, and embedded devices. The engineering work focuses on model selection, efficient indexing, and careful memory budgeting. Start small: embed a subset of data, test retrieval and local generation, and iterate on quantization and index strategies.
Build incrementally, measure on-device, and keep security at the center: when done right, local-first RAG delivers strong privacy guarantees and responsive UX without cloud dependencies.