SipLLM docs
v0.4.0 · portal v0.5 GitHub ↗

Engineering journal

Every number in these articles is measured, not projected. This is the story behind SipLLM's memory-bounded execution model, told one optimization at a time: the problem we hit in the streaming context, the change we made, and the before/after we measured on an Apple M3 (CPU-only, warm cache, median-of-3). Read the Architecture book for the module-by-module reference; read these to understand why each piece exists.

The optimizations

Streaming the layers

The founding seam: WeightSource + per-layer residency + per-row dequant, so peak RSS tracks a layer's width, not the model's depth.

Running bigger than RAM

Llama-2-13B (7.87 GB Q4) generating coherent text in 317 MB — 25× smaller than the file — on a Mac with ~3 GB free.

Streaming the LM head

The non-tied output projection was a fixed RSS floor. Streaming it in 1024-row blocks cut peak RSS 43% for a 3% decode cost.

Single-pass prefill

Naive prefill re-streamed the whole model once per prompt token. Batching over positions streams each layer once — 18× fewer bytes, bit-identical output.

Predictive prefetch

An async double-buffer ring loads block L+1 while the compute thread runs block L, hiding I/O behind arithmetic. Three backends, identical logits.

Flat resident weights

Why weight RAM stays flat as models get deeper: quantized residency + fused per-row dequant hold a layer at ~4 bits/weight instead of 32.

The Q4_K kernel

Dequantizing K-quant super-blocks a row at a time, with a NEON fast path.

Parsing GGUF

Reading unmodified Hugging Face GGUF v2/v3 files with zero dependencies.

Measuring peak RSS

Why /usr/bin/time -l is the authoritative cross-runtime figure, and how the internal counter tracks it.

Onto the phone

The Flutter/Android SDK — host-verified, not yet on-device — that turns the runtime into a pocket AI computer.

Where these fit

The founding thesis is Streaming the layers; its headline payoff is Running bigger than RAM. Each article links back to the matching page in the Architecture book.