SipLLM docs
v0.4.0 · portal v0.5 GitHub ↗

Performance history

SipLLM's trajectory, told through tags and optimization waves. The discipline is CLAUDE.md Rule 1: at the end of every wave the North Star scorecard in CHANGELOG.md is refreshed from freshly measured data and the benchmark JSON is committed under bench/results/. Peak RSS is always the authoritative cross-runtime number from /usr/bin/time -l.

The baseline is the present

The scorecard below is the current measured state (2026-07-27, Apple M3, warm cache, median-of-3) — the baseline every future wave is compared against. Where older per-release RSS/TTFT/tok/s values were not captured at the time, this page says so rather than drawing a graph of numbers that were never measured.

Current North Star scorecard (baseline)

317 MB
largest runnable
Llama-2-13B (7.87 GB Q4) — 25× smaller than its weights
121 MB
tinyllama peak RSS (stream)
… 644 MB fully resident · vs llama.cpp 1356 MB
50 tok/s
tinyllama Q8 --fast resident
vs llama.cpp 57 — ~88% at 2.1× less RAM
N/A
energy / token
needs sudo powermetrics — never fabricated

Full scorecard as committed to CHANGELOG.md:

MetricMeasured value
Peak RSStinyllama 121 MB (stream) … 644 MB (resident); smollm2 54 MB … 161 MB; vs llama.cpp CPU 1356 / 546 MB → 10–11× smaller at min budget
Resident weightsFLAT ~1.5 MB across toy 4/16/32 layers; real 37.6 MB (smollm2 1-layer) / 106.5 MB (tinyllama 1-layer); pinned dial up to 143 MB / 630 MB
Decode tok/sQ8 --fast resident smollm2 62→171 · tinyllama 50 (vs llama.cpp 57); RAM-budget dial (exact fp32) smollm2 53→66 · tinyllama 11→22
TTFTsmollm2 ~0.10 s · tinyllama ~0.68 s (vs llama.cpp 0.003 / 0.021 s)
Prefill throughputsmollm2 50–67 · tinyllama 7–23 tok/s (vs llama.cpp 1680 / 238)
Expansion factor2.7× (smollm2) · 5.5× (tinyllama) — disk ÷ peak-RSS at min budget
Largest runnableLlama-2-13B (7.87 GB Q4) in 317 MB (25×); Llama-3.1-8B (4.92 GB Q4) in 204 MB (24×)
Energy / tokenN/A (needs sudo powermetrics)

Release trajectory

v0.1.0  ──▶  v0.1.1  ──▶  v0.4.0 (Developer Preview, 2026-07-27)
 first        patch        streaming thesis proven + Flutter FFI
 tag                       (Waves 6–8)

Three tags mark the line: v0.1.0 (first tag), v0.1.1 (patch), and v0.4.0 — the Developer Preview that carries the bounded-memory waves and the on-device productization work.

Per-release metrics are being backfilled

The committed measured record (bench/results/, North Star scorecard) captures the current state comprehensively, but the project did not snapshot peak-RSS / TTFT / decode numbers at each early tag — the reproducible harness itself only landed at #33 (v0.2 baseline). So there is no honest per-tag time series for v0.1.0 → v0.1.1; those historical values are being backfilled by re-running scripts/bench.sh against the tagged commits. This page will not fabricate a trend line from numbers that were never taken.

Wave milestones (v0.4.0)

The measured progress lives in the waves, not the tags. Each entry is grounded in CHANGELOG.md.

Wave 6 — --ram-budget, the RAM↔speed dial (#37)

Turned the fixed 2-buffer streaming window into a hard byte ceiling: pin as many contiguous hot layers as fit, stream the rest, resident_bytes() ≤ budget always. This is the capability no other runtime offers — a tunable continuum between bounded RSS and speed. Measured (M3, warm, ctx 512, median-of-3):

ModelBudgetPinnedDecode tok/sStreamedPeak RSS
tinyllama0 (stream)0/2211.414411 MB121 MB
tinyllama512M13/2215.56155 MB480 MB
tinyllama768M22/2222.1576 MB644 MB
smollm20 (stream)0/3053.12824 MB54 MB
smollm2256M30/3065.9113 MB161 MB

Decode up to +95% (tinyllama) / +24% (smollm2); streamed I/O −96%; logits bit-identical across every budget (pinning is a pure cache). See the memory planner.

Wave 7 — --fast int8 SDOT kernel, near-parity Q8 decode

linear() had routed every quantized weight through fp32-dequant-then-dot, including Q8_0 — for which a tested int8 SDOT kernel existed but was dead code. --fast wires it in (opt-in; the exact fp32 path stays the oracle) and the kernel was rebuilt for ILP (vector accumulator, one horizontal reduce per row) + hardware fp16 scale conversion. Demo — TinyLlama-1.1B Q8_0, M3, ctx 512, t=4, warm:

RuntimePeak RSSDecode
llama.cpp (CPU)2326 MB~57 tok/s
SipLLM --fast --ram-budget 1200M (resident)1113 MB~50 tok/s
SipLLM --fast (streaming)175 MB6.8 tok/s

2.09× less RAM at ~88% of llama.cpp's decode, or 13× less RAM streaming. Numerically equivalent — smollm2 --fast produced byte-identical greedy output for 24 tokens; first-token predictions match. See Quantization.

Wave 8 — Flutter FFI + on-device productization (M6.5, in progress)

Makes SipLLM a phone/watch citizen without touching the runtime's math: a stable C ABI (sipllm_ffi.h) wraps the C++ engine in opaque handles + POD structs + a C token callback (whose false return is the cancellation seam), a Dart SipllmRuntime that runs inference on a worker isolate and streams Stream<SipllmToken>, on-device embeddings backed by a SQLite float32 vector store, a resumable Hugging Face downloader, and phone→Wear OS transfer.

Host-verified, not yet on-device

Wave 8 is verified on host (macOS / Apple M3): C-ABI smoke test, the full Dart → isolate → FFI path, downloader (5 tests), embedding store (13 tests), engine make test green, flutter analyze clean. It is not yet runtime-verified on device — the Android APK / on-device inference (POCO X6 Pro) and Wear transfer (OnePlus Watch 2) are the remainder of M6.5. No Android/phone performance numbers are quoted anywhere; they do not exist yet.

What comes next

Per the scorecard's own "current largest bottleneck": (1) K-quant int-dot — the --fast path is Q8_0-only, so 4-bit (Q4_K) decode still uses fp32 dequant; extending int-dot there is the biggest pending RAM headline. (2) Speculative streaming to amortize weight movement in the exceeds-RAM regime, where decode is disk-bandwidth-bound (arithmetic intensity Θ(1)). See the roadmap.