Vision — SipLLM Studio
SipLLM Studio is the offline AI computer: a six-surface workspace built entirely on top of the streaming runtime. Every screen exists to make one idea legible — running a model under a hard memory budget, larger than the RAM you have. This page describes the destination; badges mark what ships today versus what is planned.
The runtime is real and measured. Most of the Studio surfaces below are product vision, not shipped code. Shipped surfaces are <span class="badge on">shipped</span>; planned ones are <span class="badge no">planned</span>. Nothing here is measured on-device yet — the Flutter/FFI layer (Wave 8) is host-verified on Apple M3, not yet runtime-verified on a phone or watch.
The guardrail
Every surface must reinforce the same thesis: streaming under a hard memory budget. That constraint is the product.
Do not add new inference architectures or model families to chase feature parity. The engine's architecture set (Llama, Mistral, Qwen2/2.5, Gemma 2, Gemma 3 text, Phi-3, Phi-2, GPT-2, Mixtral/MoE) is deliberate. Studio's job is to visualize and control bounded-memory streaming, not to become a general model zoo.
The six surfaces
Chat shipped
Streamed token-by-token conversation on a worker isolate, cancellable mid-generate. The app applies the chat template itself (chatml/llama3/zephyr/raw) — the engine runs raw text and adds no template. Host-verified via the Dart→isolate→FFI path.
Models shipped
Model management + the resumable Hugging Face downloader (multi-connection HTTP Range, sidecar resume across restarts, sha256 verify). Backed by the device/engine capability dashboard.
Bench shipped
On-device benchmark surface over the runtime's own stats (TTFT, prefill/decode tok/s, byte/count fields). Note the Dart SipllmStats is a subset of the C struct — it drops the per-phase load_s/prefill_s/decode_s seconds.
Play planned
Offline HTML5 games in a WebView — the "AI computer" earns its keep even with no model loaded and no internet.
Labs planned
Experimental features behind a toggle, kept out of the way of the shipped surfaces until they are proven.
Settings planned
The RAM budget, thread count, scheduler policy, and backend selection — the dials that make bounded streaming tangible.
The killer surfaces (all planned)
These are what make Studio the offline AI computer rather than another chat app. All are <span class="badge no">planned</span>.
AI Playground
Live dials (--ram-budget, threads, scheduler) wired to real-time graphs of peak RSS and decode tok/s — the RAM↔speed dial made visible as you drag it.
Live Runtime Visualizer
Per-token view of every layer's state: resident · pinned · streaming · loading · evicted. Watch the streaming window slide across the model as it generates — the streaming thesis rendered live.
Storage Explorer
Tap a tensor in the GGUF directory to see its dtype, byte size, file offset, and whether it is currently streamed or resident. The model's on-disk layout, browsable.
Prompt Arena
Run the same prompt across models/budgets side by side and compare output, footprint, and speed.
AI Arcade
The model plays games — a playful stress test of on-device generation that doubles as a demo.
Why a workspace, not just a chat box
The engine's most interesting properties — a flat resident footprint across depth, a hard budget you can drag, weights streaming layer by layer — are invisible in a plain chat UI. Studio's surfaces exist to show them: the Visualizer proves peak RSS tracks layer width; the Playground proves the budget is a smooth dial; the Storage Explorer proves the model never fully materializes. Each is an argument for why streaming, made interactive.
Shipped and host-verified (Apple M3, not on-device): the Chat / Models / Bench surfaces and the FFI + isolate + downloader + embedding-store path. Planned: Play, Labs, Settings, and every killer surface above. Android APK and Wear OS transfer are the remaining work of the current wave — see Roadmap and the Android journal.