SipLLM docs
v0.4.0 · portal v0.5 GitHub ↗

Benchmarks & validation

Correctness is a measurement, not a claim; footprint is a measurement, not a claim. Every number here is Apple M3, CPU-only, warm cache, median-of-3, measured 2026-07-27. Peak RSS is the authoritative cross-runtime figure from /usr/bin/time -l. All figures are committed as JSON under bench/results/.

Host, not device

These are measured on host (Apple M3). Android on-device numbers are pending — the Flutter/FFI layer (Wave 8) is host-verified only. No phone/watch figure is quoted anywhere in these docs because none has been measured yet.

North Star scorecard

The project's single measured source of truth, refreshed at the end of every optimization wave (charter Rule 1):

MetricValue
Peak RSS — tinyllama121 MB (stream) … 644 MB (fully resident)
Peak RSS — smollm254 MB (stream) … 161 MB (fully resident)
vs llama.cpp CPU1356 MB (tinyllama) / 546 MB (smollm2) → 10–11× smaller at min budget
Resident weightsFLAT ~1.5 MB across toy 4/16/32 layers; real 37.6 MB (smollm2 1-layer) / 106.5 MB (tinyllama 1-layer); pinned dial up to 143 MB (smollm2) / 630 MB (tinyllama)
Decode tok/s (Q8 --fast, resident)smollm2 62→171 · tinyllama 50 (vs llama.cpp 57)
Decode tok/s (RAM-budget dial, exact fp32)smollm2 53→66 · tinyllama 11→22
TTFTsmollm2 ~0.10 s · tinyllama ~0.68 s (vs llama.cpp 0.003 / 0.021 s)
Prefill throughputsmollm2 50–67 · tinyllama 7–23 tok/s (vs llama.cpp 1680 / 238)
Expansion factor (disk ÷ peak-RSS, min budget)2.7× (smollm2) · 5.5× (tinyllama)
Largest runnableLlama-2-13B (7.87 GB Q4) in 317 MB (25×); Llama-3.1-8B (4.92 GB Q4) in 204 MB (24×)
Energy / tokenN/A (needs sudo powermetrics; never fabricated)
Read the trade honestly

SipLLM is not uniformly faster. TTFT and prefill throughput trail llama.cpp substantially — streaming trades disk bandwidth for a bounded footprint. The headline is RAM, and the ability to run models that do not fit at all.

Bigger than RAM — the defining capability

16 GB Mac, ~3 GB free (loading these resident is impossible), --stream-lm-head --no-async --ctx 512, greedy:

ModelWeights on diskPeak RSSModel ÷ RSS
TinyLlama-1.1B Q8_01.17 GB61 MB19×
Llama-3.1-8B Q4_K_M4.92 GB204 MB24×
Llama-2-13B Q4_K_M7.87 GB317 MB25×

Peak RSS grows with layer width, never model depth/total size. See Bigger than RAM.

Half the RAM, comparable speed

TinyLlama-1.1B Q8_0, --ctx 512, greedy, 4 threads, warm cache:

RuntimePeak RSSDecode
llama.cpp (CPU)2326 MB~57 tok/s
SipLLM --fast --ram-budget 1200M (resident)1113 MB~50 tok/s
SipLLM --fast (streaming)175 MB6.8 tok/s

2.1× less RAM at ~88% of llama.cpp's decode, or 13× less RAM streaming — a smooth RAM↔speed dial, not a fixed point.

Golden validation vs llama.cpp

For the same model and prompt, SipLLM dumps every transformer block's residual stream and the final logits and diffs them against llama.cpp's own values (captured through its eval callback). Cross-engine outputs are never bit-exact — summation order and rounding differ — so the comparison is numerical: per-layer max|Δ|, cosine similarity, and argmax/top-k agreement.

Prompt "The capital of France is" → both engines greedily predict " Paris" (TinyLlama-1.1B, 22 layers, dim 2048, GQA 32/4):

Formatworst layer max\Δ\final logit max\Δ\final cosinetop-10argmaxpeak RSSresult
F164.45e-035.76e-031.00000010/10✓412 MBPASS
Q8_01.79e-013.03e-010.9999258/10✓269 MBPASS
Q5_K_M2.37e-014.40e-010.99982910/10✓223 MBPASS
Q4_K_M3.88e-014.35e-010.99982310/10✓215 MBPASS

Two things fall straight out, both exactly what theory predicts:

  1. F16 is numerically identical to llama.cpp (cosine 1.000000) — the compute graph is correct; every residual difference in the quantized rows is pure quantization error.
  2. Error grows monotonically as quantization coarsens (F16 ≪ Q8_0 < Q5_K_M < Q4_K_M). A bug would produce erratic, layer-localized divergence; this smooth accumulation is the signature of a faithful implementation.

The golden matrix currently covers the Llama path; newer architectures are unit-tested and validated against llama.cpp as models are added. See Quantization and the LM-head journal.

RAM-budget sweep

bench_ram_budget.sh, M3, warm, ctx 512:

ModelBudgetPinnedDecode tok/sStreamedPeak RSS
tinyllama0 (stream)0/2211.414411 MB121 MB
tinyllama512M13/2215.56155 MB480 MB
tinyllama768M22/2222.1576 MB644 MB
smollm20 (stream)0/3053.12824 MB54 MB
smollm2256M30/3065.9113 MB161 MB

Pinning is a pure cache: logits + KV are bit-identical across every budget, and peak RSS ≤ budget at every point. See Memory planner.

Streamed LM head

TinyLlama-1.1B Q4_K_M, M3, warm, t=4:

ModePeak RSSResident weightsDecode
default (resident)121 MB106.5 MB11.48 tok/s
--stream-lm-head69 MB52.7 MB11.11 tok/s

Streaming the output projection nearly halves peak RSS for a negligible decode cost — the head is often the single largest tensor. See the LM-head journal.

Methodology

  • Peak RSS is read from /usr/bin/time -l (maximum resident set size) — the authoritative, cross-runtime figure, independent of any engine's self-report.
  • Median-of-3 for every timing (TTFT, decode, prefill tok/s); warm cache so the file is in the page cache.
  • Cross-engine correctness is numerical, not bit-exact — see the golden matrix above and golden/README.md for why (summation order + rounding differ between engines).
  • Every report states model · quantization · RAM budget · hardware · compiler · commit SHA alongside the metrics, and compares against the previous baseline (charter Rule 1).

Reproduce it

# peak RSS / TTFT / decode tok/s harness
scripts/bench.sh

# decode-tok/s + peak-RSS vs --ram-budget sweep
scripts/bench_ram_budget.sh

# cross-engine numerical validation vs llama.cpp
python3 golden/validate_matrix.py --prompt "The capital of France is"

# the dependency-free unit suite
make test

Latest runs are committed under bench/results/ (demo-v0.4-*.json, bigger-than-ram-*.json, ram-budget-*.json) with the full test log in bench/results/test-results-2026-07-27.txt. See the performance history for the wave-by-wave trend.