What is SipLLM?
SipLLM is a dependency-free, CPU-first LLM inference engine in C++17 that runs GGUF models larger than available RAM at a bounded, configurable peak memory ceiling. It sips weights off disk one transformer block at a time, so peak RSS tracks a single resident layer — not the whole model. It is not a llama.cpp replacement; it is a different execution model.
The one idea
Almost every runtime loads the whole model into memory before it can generate a token. A 1.1 B model in Q8_0 is ~1.1 GB resident; an 8 B model will not fit on a phone at all. SipLLM inverts that: it reads one transformer block's weights from disk with pread, runs attention + FFN, releases the block, and reuses that memory for the next one. Only a single layer's weights (plus the KV cache) are ever resident.
"LLM inference should be bounded by a configurable working set rather than total model size." Peak memory is bounded by layer width, not model depth — a deeper model of the same width has flat peak RSS.
Measured on a 16 GB Mac with only ~3 GB free (where loading these resident is impossible), --stream-lm-head --no-async --ctx 512, greedy:
What it is
- A streaming execution model. The transformer never opens a file; it talks to a
WeightSourceseam (include/llm/weight_source.h) that exposes a tensor directory plus positional reads.LayerLoader(src/loader.cpp) drives blocks through that seam with a synchronouspreadpath, an async double-buffered prefetcher, or anmmapbackend — all producing identical output. See Streaming layer loader. - A hard RAM↔speed dial.
--ram-budget Npins as many hot layers resident as fit under a byte ceiling and streams the rest, so peak weight RSS never exceeds the budget. Output is bit-for-bit identical at any budget — the dial only changes where the bytes live. See Memory planner. - A real GGUF engine. It loads unmodified GGUF v2/v3 files from Hugging Face and dispatches on
general.architecture: Llama, Mistral, Qwen2/2.5, Gemma 2, Gemma 3 text, Phi-3, Phi-2, GPT-2, and Mixtral/MoE. Dequant covers F32, F16, BF16, Q4_0/1, Q5_0/1, Q8_0 and the K-quants Q2_K, Q3_K, Q4_K, Q5_K, Q6_K, plus IQ4_NL. - Correct. For the same model and prompt, SipLLM dumps every block's residual stream and the final logits and diffs them numerically against llama.cpp. F16 is cosine
1.000000; quantization error grows monotonically as the format coarsens — the signature of a faithful implementation. See Benchmarks & validation. - Edge-first, so CPU-first. It targets phones, SBCs, and other hardware with lots of storage but little RAM/VRAM. Hand-written ARM64 NEON kernels, scalar fallbacks elsewhere.
What it is not
Not a llama.cpp replacement
It exists to prove a different execution model, and every number is measured against llama.cpp on the same hardware. Where the model fits in RAM, mature resident runtimes are faster.
Not a GPU runtime
The Vulkan backend is experimental — make VULKAN=1 enables device detection only; vulkan_matmul always falls back to CPU. Never treat it as working GPU acceleration.
Not a chat framework
The engine runs inference over raw text and adds BOS only on a fresh sequence — it applies no chat/prompt template. The Flutter app formats conversations itself (chatml/llama3/zephyr/raw).
Not fast in the extreme
Streaming an off-cache model is disk-bound (well under 1 tok/s at budget 0). That is the price of running a model that otherwise would not run at all; the RAM budget buys the speed back.
Where to go next
The argument is in Why streaming?; the honest positioning against other runtimes is in Why another runtime?; the mmap comparison is in Why not just mmap?. For the mechanics, the Architecture book is a module-by-module tour. For the roadmap and the offline-AI-computer vision, see Vision and Roadmap.