This page is the public support contract for Izwi’s current runtime surfaces. It answers four questions:
  1. Which OS and hardware combinations are supported?
  2. Which shipped artifact types expose which backends?
  3. Which deployment targets are considered supported?
  4. Which API surfaces are stable vs preview?
If another page says something different, this page should win.

Backend Matrix


Deployment Matrix


API Surface Maturity

The runtime exposes both compatibility APIs and first-party local workflow APIs under /v1. When the server is running, open /docs for the local Scalar API reference or /openapi.json for the raw OpenAPI document. The generated OpenAPI document covers the stable OpenAI-compatible contract, /v1/responses preview routes, readiness probes, and Scalar sidebar entries for preview first-party, operator, and realtime route families. Detailed preview behavior is documented in the API Reference.

CUDA Caveats

  • Linux and Windows GitHub Releases keep public binary names unchanged: izwi and izwi-server on Linux, izwi.exe and izwi-server.exe on Windows.
  • Linux and Windows GitHub Release artifacts are CPU-only and must not contain CUDA runtime libraries or private CUDA binaries.
  • Release installers do not replace the host NVIDIA driver. CUDA acceleration requires a compatible NVIDIA driver and CUDA-capable GPU.
  • Source builds still require the CUDA toolkit and remain useful for development or fallback validation.
  • The Docker CUDA image/profile is the CUDA distribution path for NVIDIA Linux hosts and may require CUDA_COMPUTE_CAP when built on a machine without nvidia-smi.
  • On macOS 15+, the recommended GPU path is Metal, not CUDA. macOS 12-14 are CPU-only.

Qwen3.8 CUDA weight residency

Qwen3.8-27B-FP8 keeps its official block-scaled FP8 checkpoint as the source artifact, but the current CUDA path does not execute native FP8 matrix multiplication. It applies each projection’s weight_scale_inv, requantizes the projection to resident Candle Q8_0 weights, and retains source-dense tensors as BF16. Runtime diagnostics identify this as resident_representation: q8_0_requantized_projections_with_dense_bf16 and fp8_execution_mode: q8_0_compressed_fallback. This compressed fallback is designed to fit 40/48 GB-class CUDA devices with a resource-fitted context, subject to actual free VRAM and allocator headroom. It must not be described as native FP8 execution or as CUDA-certified until the corresponding NVIDIA evidence is retained. The earlier fully expanded BF16 weight plan required an 80 GB-class device; CPU and Metal continue to use their expanded F32 and F16 representations respectively.

Qwen3.8 multi-token prediction defaults

Native Qwen3.8 MTP is enabled by default with one draft token on CPU, Metal, and CUDA. No environment variables are required for that default. Set IZWI_QWEN38_MTP=0 and restart or reload the model to disable MTP. The IZWI_QWEN38_MTP_DRAFT_TOKENS setting controls proposal depth rather than enablement; its default is 1, and supported explicit values are 1 through 3.

Managed Inference-State Support

This matrix describes ownership and provider classification in the current source tree. It is narrower than general model availability: a model can be available while a particular cache provider or hardware cell remains uncertified. The runtime fails model loading when a required ABI-v2 operation set or exact model/capability route is incomplete. Source/build eligibility never upgrades a route to runtime-validated or performance-certified evidence. The runtime does not silently switch back to a model-owned cache.

State topology by model family

The load path currently publishes ABI-v2 state as follows: “ABI-v2 physical ownership” is not a claim of complete model-quality or hardware performance certification. The exact model revision, capability, backend, dtype, page geometry, attention semantics, and build feature cell must pass its release lane.

CUDA context policy

CUDA uses the native context reported by the loaded chat checkpoint. CPU and Metal retain the configured sequence limit (4,096 tokens by default). The current native CUDA ceilings are: Optional YaRN extensions are not included: the local adapters do not implement their scaling parameters. Qwen3.8 fits its effective context to remaining resources without changing the 262,144-token logical maximum. Other CUDA chat routes reserve their selected paged KV capacity before publishing a model as ready. Explicit context requests still fail when the requested state cannot fit; the runtime does not overcommit VRAM. For audio generation, automatic output budgets remain bounded. Explicit CUDA requests can use the 32,768-token LFM2.5 Audio and Fish S2 contexts; Voxtral TTS accepts the upstream deployment ceiling of 2,048 output frames. CPU and Metal retain their existing limits. VibeVoice ASR uses its documented 60-minute single-pass window on CUDA by default; IZWI_VIBEVOICE_ASR_CUDA_MAX_AUDIO_SECS can lower that window but cannot raise it above 60 minutes. See the LFM2.5 Audio model card, Fish S2 config, Voxtral deployment, and VibeVoice ASR model card. Qwen3 ASR accepts up to 1,200 seconds per CUDA transcription chunk, while the forced-aligner route accepts up to 180 seconds and rejects longer inputs explicitly. CPU and Metal retain the checkpoint preprocessor window. Neither CUDA route silently truncates audio at the portable preprocessor limit. Voxtral Realtime CUDA offline decoding uses the loaded model position range (131,072 tokens in the current checkpoint) instead of the portable 1,024-frame service limit. Its physical paged cache rotates the 8,192-token attention window with one spare page; inputs beyond the loaded position range fail explicitly instead of silently dropping the tail. Architectural processing windows are not expanded. Whisper still processes 30-second encoder windows, Kokoro still uses at most 510 phonemes per model chunk, and streaming ASR/diarization families retain their bounded working buffers while the runtime orchestrates longer inputs.

Cache policy support

Eligible CUDA routes automatically use resident block-table metadata, admission-grown physical backing, device/shape-keyed attention selection, bounded sampling readback, VRAM-tiered continuous batches, and stable one-pass decode graph buckets. Graph keys include exact K/V and metadata tensor identities, geometry, dtype, scale, and softcap. Partitioned decode, FP8 storage, unobserved/pre-Ampere devices, and any capture/replay failure retain eager native decode. CPU and Metal do not enter this policy. For configuration, counters, benchmarks, and rollback, see KV Cache Operations.

Verification Guidance

Use the following expectations when validating a host:
  • Apple Silicon, macOS 15+: build or install a Metal-capable binary and run with --backend metal or IZWI_BACKEND=metal.
  • macOS 12-14: run with --backend cpu; auto and explicit metal requests fall back to CPU.
  • Linux/Windows GitHub Release: run izwi serve --backend cpu, then izwi status --detailed.
  • Docker CUDA on NVIDIA Linux hosts: run docker compose --profile cuda up, then confirm the container selects CUDA through /v1/health or izwi status --detailed from a matching client environment. Eligible Candle FlashAttention, Qwen3/Qwen3.5 RoPE, and Gemma RMSNorm CUDA routes activate automatically; their environment variables remain explicit rollback switches. Qwen3.8 is an independently gated family and must report its own execution path and evidence.
  • Linux/Windows source build for CUDA: build with cargo build --release --features cuda, then run with --backend cuda or IZWI_BACKEND=cuda. The cuda wrapper includes Candle FlashAttention, while cudnn additionally enables matching Candle/cuDNN convolution paths.

See Also