SipLLM docs
v0.4.0 · portal v0.5 GitHub ↗

Architecture book

How to read this section

This is the map. SipLLM is a layered stack with one load-bearing seam — the WeightSource — that lets the transformer run identical math whether weights come from a real GGUF, a toy .llmw, resident RAM, or streamed off disk. Each card below is a chapter; start with the streaming layer loader if you want the central idea first.

The stack

Bottom to top, every layer talks only to the one beneath it. The seam is WeightSource (include/llm/weight_source.h): a tensor directory, typed metadata, and positional read_raw/read_raw_at reads. Swapping the loader, the residency strategy, or the on-disk format touches no line of transformer math.

                    ┌──────────────────────────────────────────┐
   feeds config &   │  Runtime  (open_model → cfg → generate)   │  src/runtime.cpp
   tokens into ───▶ │  KV cache · sampler · stats               │
                    └────────────────────┬─────────────────────┘
                                         │
                    ┌────────────────────▼─────────────────────┐
                    │  Transformer  (per-block forward pass)    │  src/transformer.cpp
                    └────────────────────┬─────────────────────┘
                                         │  loadLayer(L) / unloadLayer()
                    ┌────────────────────▼─────────────────────┐
                    │  LayerLoader  (slots · roles · residency  │  src/loader.cpp
                    │  · async prefetch ring)                   │
                    └────────────────────┬─────────────────────┘
                                         │  read_raw / read_raw_at (pread or mmap-memcpy)
                    ┌────────────────────▼─────────────────────┐
                    │  WeightSource seam                        │  include/llm/weight_source.h
                    │    GgufFile (.gguf)  ·  ModelFile (.llmw)  │  src/gguf.cpp · src/format.cpp
                    └───────────────────────────────────────────┘

   underneath everything:  ThreadPool · Quant/dequant kernels · SIMD (NEON/SDOT)
   off to the side:        Device profile + Auto-tuner (thread count & scheduler)

GGUF parsing and the Tokenizer feed this stack from the left — GGUF metadata becomes the ModelConfig and the tokenizer vocab; both are built once through the same WeightSource. The compute floor (ThreadPool, quant kernels, SIMD) and the device/auto-tune path sit underneath as shared services.

The forward pass in one line

embed_token → for L in 0..n_layers: loadLayer(L) · RMSNorm · QKV · RoPE · GQA attention · out-proj · RMSNorm · SwiGLU · unloadLayer() → final RMSNorm · project_output → logits — and only one block's weights are resident at a time unless you pin more with --ram-budget.

Chapters

Streaming layer loader

The seam that makes memory track layer size, not model size. Read →

Layer residency

Pinned vs streamed layers and the quantized-vs-fp32 residency modes. Read →

Memory planner

The --ram-budget contract planner that pins hot layers under a hard peak-RSS ceiling. Read →

Predictive prefetch

The async double-buffered ring that materializes block L+1 while block L computes. Read →

Runtime

Owns the whole stack: open_model → ModelConfig → KV cache → sampler → generate loop. Read →

Transformer

The block math: RMSNorm, GQA attention, RoPE, SwiGLU, per-arch dispatch across nine architectures. Read →

KV cache

Grow-on-demand key/value storage sized to the actual sequence, not the context ceiling. Read →

Quantization

Per-row dequantize-then-dot, K-quant NEON fast paths, and the Q8_0 int8 SDOT --fast kernel. Read →

GGUF parser

The real GGUF v2/v3 reader (and its .llmw sibling) behind the WeightSource seam. Read →

Tokenizer

One tokenizer, three kinds — Byte, SentencePiece, byte-level BPE — all built from GGUF metadata. Read →

Thread pool

The pthread work-sharing pool and its schedule policies that parallelize every matmul. Read →

Auto-tuning

Micro-benchmarks that pick thread count and scheduler per machine, cached under ~/.sipllm. Read →

Flutter runtime

The FFI bridge and Dart SDK that turn the engine into an offline on-device runtime. Read →

Wear support

Bringing the streaming runtime down to a watch-class device. Read →

Where to go next

New to the thesis? Read Why streaming? and the streaming-layers journal. Want the measured payoff? See Benchmarks. Building on the engine? Jump to the API reference.