SipLLM docs
v0.4.0 · portal v0.5 GitHub ↗

Streaming layer loader

Architecture book

Problem → Design → Alternatives → Tradeoffs → Source → Diagrams → Future. The loader is the seam that makes the memory-bounded thesis true. For the reference of every function named here, see the C++ engine API.

Problem

An LLM's weights dwarf its activations. Loading them all resident makes peak RSS scale with model size — 1.1 GB for a 1.1 B Q8 model, tens of GB for a 70 B one, which simply does not fit on a phone. But a decoder only touches one transformer block at a time: block L's weights are needed only while block L runs. If we could make exactly one block resident and reuse that memory for the next, peak weight RSS would track layer size (model width), not model size (depth). The loader is what makes that reuse safe, fast, and invisible to the math.

Design

The transformer never opens a file. It talks to a WeightSource (include/llm/weight_source.h) that exposes three things: a tensor directory (tensors(), find()), typed metadata (meta_int/meta_float/meta_str), and positional reads (read_raw() / read_raw_at(), a pread that never loads the whole file). Both the real GGUF parser (GgufFile) and the toy .llmw reader (ModelFile) implement it, so swapping loaders or residency strategy touches no line of math.

LayerLoader (include/llm/loader.h, src/loader.cpp) drives blocks through that seam. Each block asks for its weights by Role (28 roles; the core nine are AttnNorm, AttnQ, AttnK, AttnV, AttnOut, FfnNorm, FfnGate, FfnUp, FfnDown), which the loader maps to GGUF tensor names like blk.<L>.attn_q.weight. The forward pass loops:

embed_token(token)              -> x   (one row streamed, or resident if tied)
for layer L in 0 .. n_layers-1:
    loadLayer(L)                -> make block L's weights resident (may block on prefetch)
    block(L, pos)               -> RMSNorm -> QKV -> RoPE -> GQA attn -> proj -> RMSNorm -> SwiGLU
    unloadLayer()               -> release block L for reuse
final RMSNorm + project_output  -> logits

Two mechanisms keep only a layer's worth of weights live:

  1. Per-layer residency. loadLayer(L) reads block L into a slot buffer; the slot is recycled for a later layer. With the synchronous single-buffer path exactly one block is resident.
  2. Weights stay quantized; dequant is per-row. In Residency::Quantized the loader keeps each weight's raw on-disk bytes resident and never bulk-expands them. matmul_quant dequantizes one output row into a tiny scratch buffer, dots it with the activation, and moves on — so a whole layer stays ~4 bits/weight in RAM instead of 32. Residency::FP32 (dequant on load) exists mainly as the numeric oracle. Norm weights are tiny and always fp32.
on disk — GGUF, quantized (may be tens of GB)block 0block 1block 2…block N-1out headpread one blockresident in RAM — bounded1–2 layer buffersattn + FFNKV cachepeak RAM ≈one layer + KVfree & reuse the buffer for the next block

There are three interchangeable backends, all producing identical output:

  • Synchronous single-buffer pread — one slot, one block resident (--no-async).
  • Async double-buffered prefetch ring (default) — a background worker materializes block L+1 while the compute thread runs block L. See Predictive prefetch.
  • mmap — read_raw_at copies from mapped pages instead of pread, letting the OS page cache buffer (--mmap).

Alternatives considered

ApproachPeak RSSWhy not the default
Load everything resident (llama.cpp default)∝ model sizeCannot run a model larger than RAM at all — the entire point.
mmap the whole file∝ working set, uncontrolledNo hard ceiling; the OS decides residency and can still fault the whole model in. SipLLM offers mmap as one backend, but pairs it with an explicit budget.
Dequantize the layer to fp32 on load8× the quantized layerWastes the dominant memory term; kept only as Residency::FP32 for the numeric oracle.
Keep every layer, evict LRUtunable but complexThe memory planner + layer residency achieve the same RAM↔speed control with a simpler contiguous-pin model.

Tradeoffs

Streaming trades disk bandwidth for memory. In the bounded-memory extreme (budget 0) decode is disk-bound — arithmetic intensity is Θ(1), so an off-cache model runs well under 1 tok/s. That is the price of running a model that otherwise would not run at all. The --ram-budget dial buys the speed back linearly: pin more hot layers, re-stream fewer per token. Correctness is never traded — every backend and every budget yields bit-identical logits, guarded by tests/test_e2e.cpp and tests/test_ram_budget.cpp.

Source files

FileRole
include/llm/weight_source.hthe WeightSource seam (directory + metadata + positional reads)
include/llm/loader.h / src/loader.cppLayerLoader: slots, roles, residency, the prefetch worker
include/llm/file_backing.hPOSIX pread_exact / optional mmap shared by every source
src/quant.cppmatmul_quant — the per-row dequantize-then-dot
src/transformer.cppthe forward pass that drives loadLayer/unloadLayer

Future work

  • Deeper prefetch pipelining and explicit posix_fadvise/madvise(WILLNEED) readahead (FileBacking::prefetch() exists but is currently unused).
  • Non-contiguous / hotness-aware pinning — today pinning always covers the leading run [0, n_pinned); there is no hotness heuristic.
  • Speculative streaming to amortize weight movement in the exceeds-RAM regime (the long-term throughput moat).