Contributing
SipLLM is a small, readable, dependency-free codebase — genuinely nice to hack on. This page distills CONTRIBUTING.md and the engineering charter (CLAUDE.md) into what you need to land a change without regressing the thesis.
No new runtime dependencies. The whole point is a from-scratch engine in standard C++17 + pthread — no PyTorch, no ONNX, no ggml, no BLAS, no CMake. A PR that adds a third-party library to the inference path will not be merged. (Dev-only tooling is negotiable.)
Build
git clone https://github.com/ankit1057/sipllm.git && cd sipllm
make # -> build/llm, build/bench, build/inspect_gguf, ...
make test # 34 unit tests, no external deps — should be all green
That is the entire toolchain: make + a C++17 compiler. If make test is green, you have a working dev environment.
Coding standards
- C++17, zero runtime deps — the hard rule above.
- Match the surrounding style — ~100-col lines; comments explain the why
- Additive & guarded — a new dial's default must reproduce prior behavior
Measure-first / North Star discipline
CLAUDE.md Rule 0: no optimization begins without measurement. Before touching an issue, run the benchmark suite (scripts/bench.sh, scripts/bench_ram_budget.sh), rank bottlenecks by impact × leverage, and state the current bottleneck · expected benefit · risk · complexity · how success is measured · rollback plan. Every change must earn its place against the eight levers (reduce peak RSS, raise throughput, cut latency, grow the largest runnable model, improve correctness/portability, simplify, improve maintainability).
Rule 1: at the end of every optimization wave, refresh the North Star scorecard at the top of CHANGELOG.md with freshly measured values, append a dated wave entry, and commit the latest benchmark JSON under bench/results/. The scorecard is the single measured source of truth.
Regression policy
- The golden matrix must stay green. If you touch the forward pass,
python3 golden/validate_matrix.py --prompt "The capital of France is"
make teststays green on every change.--ram-budget 0is byte-identical to prior behavior — pinning is a pure
Benchmark expectations
Every benchmark report states model · quantization · RAM budget · hardware · compiler · commit SHA, alongside peak RSS · resident weights · decode tok/s · prefill tok/s · TTFT, and compares against the previous baseline. Numbers are authoritative only from /usr/bin/time -l (peak RSS, the cross-runtime figure) and the engine's own stats. Never fabricate a metric — mark it N/A if unmeasured (e.g. energy/token needs sudo powermetrics).
PR checklist
- Fork, branch from
main(git checkout -b feature/my-thing). - Make the change; add or adjust tests.
make testgreen — and the golden matrix if you touched the math.- If it is an optimization: attach before/after measurements (full report
- Confirm additive & guarded — defaults reproduce prior behavior byte-for-byte.
- Open the PR with a clear what and why; reference any issue / RFC.
- CI (build + tests on gcc and clang) must pass.
Testing philosophy
Zero-dep harness
34 unit tests, no gtest/catch2 — the test suite honors the same no-dependency rule as the engine (make test).
ref_forward oracle
The exact fp32-dequant path is the correctness oracle. Approximations like --fast (int8 SDOT) are opt-in and validated against it, never the other way round.
Two regimes
Bit-identical where a change must not move a bit (--ram-budget pinning, KV grow-on-demand); tolerance (per-layer cosine ≈ 1.0, 1e-3 max|Δ|) for cross-engine quantized comparison.
Sanitizer / valgrind gates
CI gates on ASan / UBSan / TSan + valgrind + a prompt/config fuzzer — the whole uninitialized-read class was eliminated at source (#32) and is kept out by these gates.
Good first issues
- New quantization format — a dequant path in
src/quant.cpp+ a round-trip - More registry models — extend
builtin_url()insipllmwith public, - NEON / x86 SIMD — port a scalar hot path in
src/neon.cpp(guarded by - Docs — a diagram or a forward-pass walkthrough.
By contributing you agree your work is licensed under the project's MIT License. For the reasoning behind these rules, see design decisions and the RFC index.