SipLLM docs
v0.4.0 · portal v0.5 GitHub ↗

SipLLM

Run GGUF models larger than available RAM through bounded-memory transformer layer streaming. SipLLM is a dependency-free, CPU-first LLM inference engine in C++17 that sips weights off disk one transformer block at a time, so peak memory tracks a single resident layer — not the whole model. A Flutter/Android SDK turns that runtime into an offline AI computer you carry in your pocket.

Start anywhere:

The one idea

The usual way to run an LLM loads the entire model into memory. A 1.1 B model in Q8_0 is ~1.1 GB resident; an 8 B model will not fit on a phone at all. SipLLM instead streams one transformer block at a time: pread the block's weights from disk, run attention + FFN, free it, move on. Only a single layer's weights (plus the KV cache) are ever resident, so a model many times larger than RAM still runs — memory is bounded by layer size, not model size.

Central hypothesis

"LLM inference should be bounded by a configurable working set rather than total model size." SipLLM is not trying to replace llama.cpp — it exists to prove a different execution model, and every number here is measured against llama.cpp on the same hardware.

Measured, not claimed

Apple M3, CPU-only, warm cache, median-of-3; peak RSS from /usr/bin/time -l (the authoritative cross-runtime figure). Reproducible and committed as JSON under bench/results/.

317 MB
peak RSS to run Llama-2-13B
7.87 GB Q4 weights → 25× smaller than the model
2.1×
less RAM than llama.cpp
at ~88% of its decode (TinyLlama Q8, resident)
13×
less RAM, streaming
same TinyLlama in 175 MB vs 2326 MB
0
runtime dependencies
standard C++17 + pthreads only

Bigger than RAM — the defining capability

Measured on a 16 GB Mac with only ~3 GB free, where loading these models resident is impossible (--stream-lm-head --no-async --ctx 512, greedy):

ModelWeights on diskPeak RSSModel ÷ RSSOutput
TinyLlama-1.1B (Q8_0)1.17 GB61 MB19×coherent
Llama-3.1-8B (Q4_K_M)4.92 GB204 MB24×coherent
Llama-2-13B (Q4_K_M)7.87 GB317 MB25×coherent

Peak RSS grows with layer width, never model depth/total size — a deeper model of the same width has flat peak RSS.

Half the RAM, comparable speed

TinyLlama-1.1B Q8_0, --ctx 512, greedy, 4 threads, warm cache:

RuntimePeak RSSDecodevs llama.cpp
llama.cpp (CPU, -ngl 0 -t 4)2326 MB~57 tok/sbaseline
SipLLM --fast --ram-budget 1200M (resident)1113 MB~50 tok/s2.1× less RAM, ~12% slower
SipLLM --fast (streaming)175 MB6.8 tok/s13.3× less RAM

It is a smooth RAM↔speed dial, not a fixed point. Output is numerically equivalent — see Benchmarks & validation for the layer-by-layer diff against llama.cpp.

What is inside

Real GGUF parser

Loads unmodified GGUF v2/v3 files from Hugging Face — metadata, tensor directory, and the common quantizations.

Streaming layer loader

Synchronous pread, an async double-buffered prefetcher, or an mmap backend — switchable and benchmarked side by side.

RAM↔speed dial

--ram-budget pins as many hot layers as fit under a hard peak-RSS ceiling and streams the rest. Output is bit-identical at any budget.

Nine architectures

Llama, Mistral, Qwen2/2.5, Gemma 2, Gemma 3 text, Phi-3, Phi-2, GPT-2, and Mixtral/MoE — dispatched on general.architecture.

ARM64 NEON kernels

sdot-accelerated int8 matmul with an x86 AVX2/FMA path and scalar fallbacks everywhere else.

Flutter / Android SDK

A stable C ABI wraps the engine; a Dart isolate streams tokens to a phone & Wear OS UI with downloads, embeddings, and benchmarks — all offline.

Quick start

# one-liner installer (prebuilt release, or builds from source)
curl -fsSL https://raw.githubusercontent.com/ankit1057/sipllm/main/install.sh | sh
sipllm run tinyllama -p "The capital of France is"

# from source — the entire toolchain is make + a C++17 compiler
git clone https://github.com/ankit1057/sipllm.git && cd sipllm
make            # -> build/llm, build/bench, build/inspect_gguf, ...
make test       # dependency-free unit suite, all green

# run a model bigger than your RAM
./build/llm model.gguf -p "prompt" -n 40 --stream-lm-head --no-async
./build/llm model.gguf -p "prompt" --ram-budget 512M   # cap peak RSS
Where to go next

New here? Read What is SipLLM? and Why streaming?. Want the mechanics? The Architecture book is a module-by-module tour, and the Engineering journal tells the story behind each optimization. Building on it? Jump to the API reference.