Engineering journal
Every number in these articles is measured, not projected. This is the story behind SipLLM's memory-bounded execution model, told one optimization at a time: the problem we hit in the streaming context, the change we made, and the before/after we measured on an Apple M3 (CPU-only, warm cache, median-of-3). Read the Architecture book for the module-by-module reference; read these to understand why each piece exists.
The optimizations
Streaming the layers
The founding seam: WeightSource + per-layer residency + per-row dequant, so peak RSS tracks a layer's width, not the model's depth.
Running bigger than RAM
Llama-2-13B (7.87 GB Q4) generating coherent text in 317 MB — 25× smaller than the file — on a Mac with ~3 GB free.
Streaming the LM head
The non-tied output projection was a fixed RSS floor. Streaming it in 1024-row blocks cut peak RSS 43% for a 3% decode cost.
Single-pass prefill
Naive prefill re-streamed the whole model once per prompt token. Batching over positions streams each layer once — 18× fewer bytes, bit-identical output.
Predictive prefetch
An async double-buffer ring loads block L+1 while the compute thread runs block L, hiding I/O behind arithmetic. Three backends, identical logits.
Flat resident weights
Why weight RAM stays flat as models get deeper: quantized residency + fused per-row dequant hold a layer at ~4 bits/weight instead of 32.
The Q4_K kernel
Dequantizing K-quant super-blocks a row at a time, with a NEON fast path.
Parsing GGUF
Reading unmodified Hugging Face GGUF v2/v3 files with zero dependencies.
Measuring peak RSS
Why /usr/bin/time -l is the authoritative cross-runtime figure, and how the internal counter tracks it.
Onto the phone
The Flutter/Android SDK — host-verified, not yet on-device — that turns the runtime into a pocket AI computer.
The founding thesis is Streaming the layers; its headline payoff is Running bigger than RAM. Each article links back to the matching page in the Architecture book.