SipLLM
Run GGUF models larger than available RAM through bounded-memory transformer layer streaming. SipLLM is a dependency-free, CPU-first LLM inference engine in C++17 that sips weights off disk one transformer block at a time, so peak memory tracks a single resident layer — not the whole model. A Flutter/Android SDK turns that runtime into an offline AI computer you carry in your pocket.
Start anywhere:
The one idea
The usual way to run an LLM loads the entire model into memory. A 1.1 B model in Q8_0 is ~1.1 GB resident; an 8 B model will not fit on a phone at all. SipLLM instead streams one transformer block at a time: pread the block's weights from disk, run attention + FFN, free it, move on. Only a single layer's weights (plus the KV cache) are ever resident, so a model many times larger than RAM still runs — memory is bounded by layer size, not model size.
"LLM inference should be bounded by a configurable working set rather than total model size." SipLLM is not trying to replace llama.cpp — it exists to prove a different execution model, and every number here is measured against llama.cpp on the same hardware.
Measured, not claimed
Apple M3, CPU-only, warm cache, median-of-3; peak RSS from /usr/bin/time -l (the authoritative cross-runtime figure). Reproducible and committed as JSON under bench/results/.
Bigger than RAM — the defining capability
Measured on a 16 GB Mac with only ~3 GB free, where loading these models resident is impossible (--stream-lm-head --no-async --ctx 512, greedy):
| Model | Weights on disk | Peak RSS | Model ÷ RSS | Output |
|---|---|---|---|---|
| TinyLlama-1.1B (Q8_0) | 1.17 GB | 61 MB | 19× | coherent |
| Llama-3.1-8B (Q4_K_M) | 4.92 GB | 204 MB | 24× | coherent |
| Llama-2-13B (Q4_K_M) | 7.87 GB | 317 MB | 25× | coherent |
Peak RSS grows with layer width, never model depth/total size — a deeper model of the same width has flat peak RSS.
Half the RAM, comparable speed
TinyLlama-1.1B Q8_0, --ctx 512, greedy, 4 threads, warm cache:
| Runtime | Peak RSS | Decode | vs llama.cpp |
|---|---|---|---|
llama.cpp (CPU, -ngl 0 -t 4) | 2326 MB | ~57 tok/s | baseline |
SipLLM --fast --ram-budget 1200M (resident) | 1113 MB | ~50 tok/s | 2.1× less RAM, ~12% slower |
SipLLM --fast (streaming) | 175 MB | 6.8 tok/s | 13.3× less RAM |
It is a smooth RAM↔speed dial, not a fixed point. Output is numerically equivalent — see Benchmarks & validation for the layer-by-layer diff against llama.cpp.
What is inside
Real GGUF parser
Loads unmodified GGUF v2/v3 files from Hugging Face — metadata, tensor directory, and the common quantizations.
Streaming layer loader
Synchronous pread, an async double-buffered prefetcher, or an mmap backend — switchable and benchmarked side by side.
RAM↔speed dial
--ram-budget pins as many hot layers as fit under a hard peak-RSS ceiling and streams the rest. Output is bit-identical at any budget.
Nine architectures
Llama, Mistral, Qwen2/2.5, Gemma 2, Gemma 3 text, Phi-3, Phi-2, GPT-2, and Mixtral/MoE — dispatched on general.architecture.
ARM64 NEON kernels
sdot-accelerated int8 matmul with an x86 AVX2/FMA path and scalar fallbacks everywhere else.
Flutter / Android SDK
A stable C ABI wraps the engine; a Dart isolate streams tokens to a phone & Wear OS UI with downloads, embeddings, and benchmarks — all offline.
Quick start
# one-liner installer (prebuilt release, or builds from source)
curl -fsSL https://raw.githubusercontent.com/ankit1057/sipllm/main/install.sh | sh
sipllm run tinyllama -p "The capital of France is"
# from source — the entire toolchain is make + a C++17 compiler
git clone https://github.com/ankit1057/sipllm.git && cd sipllm
make # -> build/llm, build/bench, build/inspect_gguf, ...
make test # dependency-free unit suite, all green
# run a model bigger than your RAM
./build/llm model.gguf -p "prompt" -n 40 --stream-lm-head --no-async
./build/llm model.gguf -p "prompt" --ram-budget 512M # cap peak RSS
New here? Read What is SipLLM? and Why streaming?. Want the mechanics? The Architecture book is a module-by-module tour, and the Engineering journal tells the story behind each optimization. Building on it? Jump to the API reference.