SipLLM docs
v0.4.0 · portal v0.5 GitHub ↗

API reference

SipLLM exposes three programming surfaces, layered strictly from bottom to top. Each is a thin wrapper over the one below it — the C++ engine does the math, a stable C ABI flattens it into FFI-safe POD, and the Dart/Flutter package drives that ABI from a background isolate.

Reference pages

These are reference pages: signatures are reproduced verbatim from source (include/llm/*.h, bindings/flutter/sipllm_flutter/ffi/sipllm_ffi.h, lib/src/**), grouped by type, one line of semantics each. For the why behind each subsystem, follow the links into the architecture book.

The three surfaces

SurfaceHeader / entry pointLanguageAudience
C++ engineinclude/llm/*.h (namespace llm)C++17Embedders linking the engine directly; the CLI (build/llm) and web server build on it.
Stable C ABIbindings/flutter/sipllm_flutter/ffi/sipllm_ffi.hCAny language with a C FFI. Opaque handles + POD structs + function-pointer callbacks.
Dart / Flutterpackage:sipllm_flutterDartFlutter apps. Isolate-backed runtime plus download / embedding / device / Wear helpers.

How they relate

The C++ engine (llm::Runtime and friends) has a rich C++17 interface — std::string, std::function, std::unique_ptr — none of which is FFI-safe. The C ABI (sipllm_ffi.h) is the only surface non-C++ callers touch: it is intentionally additive over the engine, adding no math and changing no defaults (sipllm_params zero-initialized reproduces the CLI's behavior). Dart's SipllmRuntime in turn wraps that ABI, running sipllm_generate on a worker isolate so the UI thread never blocks, and calling the thread-safe sipllm_cancel from the main isolate.

Dart / FlutterSipllmRuntime (isolate)Download · EmbeddingDevice · WearC ABIsipllm_ffi.hopaque ctx + POD18 SIPLLM_API fnsC++ engineRuntime · LayerLoaderTransformer · KVCacheops · quant · tokenizer

Pick a surface

C ABI →

Bind SipLLM from any language. Four POD structs, three enums, one callback typedef, and 18 SIPLLM_API functions. See the C ABI reference.

C++ engine →

Link namespace llm directly for the full streaming stack: open_model(), Runtime, LayerLoader, Transformer, ModelConfig, Tokenizer, and the ops/quant kernels. See the C++ engine reference.

Dart / Flutter →

Add package:sipllm_flutter: SipllmRuntime, plus a resumable Hugging Face downloader, a SQLite embedding store, device/arch detection, and phone↔Wear transfer. See the Dart reference.

Capability caveats that cut across all three surfaces
  • GPU offload is detection-only. sipllm_vulkan_available() / SipllmDevice.vulkanAvailable can report a device, but vulkan_matmul always falls back to CPU — never assume working GPU acceleration.
  • The engine applies no chat template. All three surfaces run inference over raw text (BOS added only on a fresh sequence). Conversation formatting is the caller's job (the Flutter app does it in PromptTemplate).
  • embed() clears KV state on every surface — use a dedicated context/runtime for embeddings, not one mid-conversation.
  • Q8_1 and Q8_K are not dequantizable (they throw); the --fast int8 SDOT kernel is Q8_0-only and ARM-only.

See also