Streaming layer loader
Problem → Design → Alternatives → Tradeoffs → Source → Diagrams → Future. The loader is the seam that makes the memory-bounded thesis true. For the reference of every function named here, see the C++ engine API.
Problem
An LLM's weights dwarf its activations. Loading them all resident makes peak RSS scale with model size — 1.1 GB for a 1.1 B Q8 model, tens of GB for a 70 B one, which simply does not fit on a phone. But a decoder only touches one transformer block at a time: block L's weights are needed only while block L runs. If we could make exactly one block resident and reuse that memory for the next, peak weight RSS would track layer size (model width), not model size (depth). The loader is what makes that reuse safe, fast, and invisible to the math.
Design
The transformer never opens a file. It talks to a WeightSource (include/llm/weight_source.h) that exposes three things: a tensor directory (tensors(), find()), typed metadata (meta_int/meta_float/meta_str), and positional reads (read_raw() / read_raw_at(), a pread that never loads the whole file). Both the real GGUF parser (GgufFile) and the toy .llmw reader (ModelFile) implement it, so swapping loaders or residency strategy touches no line of math.
LayerLoader (include/llm/loader.h, src/loader.cpp) drives blocks through that seam. Each block asks for its weights by Role (28 roles; the core nine are AttnNorm, AttnQ, AttnK, AttnV, AttnOut, FfnNorm, FfnGate, FfnUp, FfnDown), which the loader maps to GGUF tensor names like blk.<L>.attn_q.weight. The forward pass loops:
embed_token(token) -> x (one row streamed, or resident if tied)
for layer L in 0 .. n_layers-1:
loadLayer(L) -> make block L's weights resident (may block on prefetch)
block(L, pos) -> RMSNorm -> QKV -> RoPE -> GQA attn -> proj -> RMSNorm -> SwiGLU
unloadLayer() -> release block L for reuse
final RMSNorm + project_output -> logits
Two mechanisms keep only a layer's worth of weights live:
- Per-layer residency.
loadLayer(L)reads blockLinto a slot buffer; the slot is recycled for a later layer. With the synchronous single-buffer path exactly one block is resident. - Weights stay quantized; dequant is per-row. In
Residency::Quantizedthe loader keeps each weight's raw on-disk bytes resident and never bulk-expands them.matmul_quantdequantizes one output row into a tiny scratch buffer, dots it with the activation, and moves on — so a whole layer stays ~4 bits/weight in RAM instead of 32.Residency::FP32(dequant on load) exists mainly as the numeric oracle. Norm weights are tiny and always fp32.
There are three interchangeable backends, all producing identical output:
- Synchronous single-buffer
pread— one slot, one block resident (--no-async). - Async double-buffered prefetch ring (default) — a background worker materializes block
L+1while the compute thread runs blockL. See Predictive prefetch. mmap—read_raw_atcopies from mapped pages instead ofpread, letting the OS page cache buffer (--mmap).
Alternatives considered
| Approach | Peak RSS | Why not the default |
|---|---|---|
| Load everything resident (llama.cpp default) | ∝ model size | Cannot run a model larger than RAM at all — the entire point. |
mmap the whole file | ∝ working set, uncontrolled | No hard ceiling; the OS decides residency and can still fault the whole model in. SipLLM offers mmap as one backend, but pairs it with an explicit budget. |
| Dequantize the layer to fp32 on load | 8× the quantized layer | Wastes the dominant memory term; kept only as Residency::FP32 for the numeric oracle. |
| Keep every layer, evict LRU | tunable but complex | The memory planner + layer residency achieve the same RAM↔speed control with a simpler contiguous-pin model. |
Tradeoffs
Streaming trades disk bandwidth for memory. In the bounded-memory extreme (budget 0) decode is disk-bound — arithmetic intensity is Θ(1), so an off-cache model runs well under 1 tok/s. That is the price of running a model that otherwise would not run at all. The --ram-budget dial buys the speed back linearly: pin more hot layers, re-stream fewer per token. Correctness is never traded — every backend and every budget yields bit-identical logits, guarded by tests/test_e2e.cpp and tests/test_ram_budget.cpp.
Source files
| File | Role |
|---|---|
include/llm/weight_source.h | the WeightSource seam (directory + metadata + positional reads) |
include/llm/loader.h / src/loader.cpp | LayerLoader: slots, roles, residency, the prefetch worker |
include/llm/file_backing.h | POSIX pread_exact / optional mmap shared by every source |
src/quant.cpp | matmul_quant — the per-row dequantize-then-dot |
src/transformer.cpp | the forward pass that drives loadLayer/unloadLayer |
Future work
- Deeper prefetch pipelining and explicit
posix_fadvise/madvise(WILLNEED)readahead (FileBacking::prefetch()exists but is currently unused). - Non-contiguous / hotness-aware pinning — today pinning always covers the leading run
[0, n_pinned); there is no hotness heuristic. - Speculative streaming to amortize weight movement in the exceeds-RAM regime (the long-term throughput moat).