Architecture book
This is the map. SipLLM is a layered stack with one load-bearing seam — the WeightSource — that lets the transformer run identical math whether weights come from a real GGUF, a toy .llmw, resident RAM, or streamed off disk. Each card below is a chapter; start with the streaming layer loader if you want the central idea first.
The stack
Bottom to top, every layer talks only to the one beneath it. The seam is WeightSource (include/llm/weight_source.h): a tensor directory, typed metadata, and positional read_raw/read_raw_at reads. Swapping the loader, the residency strategy, or the on-disk format touches no line of transformer math.
┌──────────────────────────────────────────┐
feeds config & │ Runtime (open_model → cfg → generate) │ src/runtime.cpp
tokens into ───▶ │ KV cache · sampler · stats │
└────────────────────┬─────────────────────┘
│
┌────────────────────▼─────────────────────┐
│ Transformer (per-block forward pass) │ src/transformer.cpp
└────────────────────┬─────────────────────┘
│ loadLayer(L) / unloadLayer()
┌────────────────────▼─────────────────────┐
│ LayerLoader (slots · roles · residency │ src/loader.cpp
│ · async prefetch ring) │
└────────────────────┬─────────────────────┘
│ read_raw / read_raw_at (pread or mmap-memcpy)
┌────────────────────▼─────────────────────┐
│ WeightSource seam │ include/llm/weight_source.h
│ GgufFile (.gguf) · ModelFile (.llmw) │ src/gguf.cpp · src/format.cpp
└───────────────────────────────────────────┘
underneath everything: ThreadPool · Quant/dequant kernels · SIMD (NEON/SDOT)
off to the side: Device profile + Auto-tuner (thread count & scheduler)
GGUF parsing and the Tokenizer feed this stack from the left — GGUF metadata becomes the ModelConfig and the tokenizer vocab; both are built once through the same WeightSource. The compute floor (ThreadPool, quant kernels, SIMD) and the device/auto-tune path sit underneath as shared services.
The forward pass in one line
embed_token → for L in 0..n_layers: loadLayer(L) · RMSNorm · QKV · RoPE · GQA attention · out-proj · RMSNorm · SwiGLU · unloadLayer() → final RMSNorm · project_output → logits — and only one block's weights are resident at a time unless you pin more with --ram-budget.
Chapters
Streaming layer loader
The seam that makes memory track layer size, not model size. Read →
Layer residency
Pinned vs streamed layers and the quantized-vs-fp32 residency modes. Read →
Memory planner
The --ram-budget contract planner that pins hot layers under a hard peak-RSS ceiling. Read →
Predictive prefetch
The async double-buffered ring that materializes block L+1 while block L computes. Read →
Runtime
Owns the whole stack: open_model → ModelConfig → KV cache → sampler → generate loop. Read →
Transformer
The block math: RMSNorm, GQA attention, RoPE, SwiGLU, per-arch dispatch across nine architectures. Read →
KV cache
Grow-on-demand key/value storage sized to the actual sequence, not the context ceiling. Read →
Quantization
Per-row dequantize-then-dot, K-quant NEON fast paths, and the Q8_0 int8 SDOT --fast kernel. Read →
GGUF parser
The real GGUF v2/v3 reader (and its .llmw sibling) behind the WeightSource seam. Read →
Tokenizer
One tokenizer, three kinds — Byte, SentencePiece, byte-level BPE — all built from GGUF metadata. Read →
Thread pool
The pthread work-sharing pool and its schedule policies that parallelize every matmul. Read →
Auto-tuning
Micro-benchmarks that pick thread count and scheduler per machine, cached under ~/.sipllm. Read →
Flutter runtime
The FFI bridge and Dart SDK that turn the engine into an offline on-device runtime. Read →
Wear support
Bringing the streaming runtime down to a watch-class device. Read →
Where to go next
New to the thesis? Read Why streaming? and the streaming-layers journal. Want the measured payoff? See Benchmarks. Building on the engine? Jump to the API reference.