- Which OS and hardware combinations are supported?
- Which shipped artifact types expose which backends?
- Which deployment targets are considered supported?
- Which API surfaces are stable vs preview?
Backend Matrix
Deployment Matrix
API Surface Maturity
The runtime exposes both compatibility APIs and first-party local workflow APIs under/v1.
When the server is running, open /docs for the local Scalar API reference or
/openapi.json for the raw OpenAPI document. The generated OpenAPI document
covers the stable OpenAI-compatible contract, /v1/responses preview routes,
readiness probes, and Scalar sidebar entries for preview first-party, operator,
and realtime route families. Detailed preview behavior is documented in the
API Reference.
CUDA Caveats
- Linux and Windows GitHub Releases keep public binary names unchanged:
izwiandizwi-serveron Linux,izwi.exeandizwi-server.exeon Windows. - Linux and Windows GitHub Release artifacts are CPU-only and must not contain CUDA runtime libraries or private CUDA binaries.
- Release installers do not replace the host NVIDIA driver. CUDA acceleration requires a compatible NVIDIA driver and CUDA-capable GPU.
- Source builds still require the CUDA toolkit and remain useful for development or fallback validation.
- The Docker CUDA image/profile is the CUDA distribution path for NVIDIA Linux hosts and may require
CUDA_COMPUTE_CAPwhen built on a machine withoutnvidia-smi. - On macOS 15+, the recommended GPU path is Metal, not CUDA. macOS 12-14 are CPU-only.
Qwen3.8 CUDA weight residency
Qwen3.8-27B-FP8 keeps its official block-scaled FP8 checkpoint as the source
artifact, but the current CUDA path does not execute native FP8 matrix
multiplication. It applies each projection’s weight_scale_inv, requantizes
the projection to resident Candle Q8_0 weights, and retains source-dense tensors
as BF16. Runtime diagnostics identify this as
resident_representation: q8_0_requantized_projections_with_dense_bf16 and
fp8_execution_mode: q8_0_compressed_fallback.
This compressed fallback is designed to fit 40/48 GB-class CUDA devices with a
resource-fitted context, subject to actual free VRAM and allocator headroom. It
must not be described as native FP8 execution or as CUDA-certified until the
corresponding NVIDIA evidence is retained. The earlier fully expanded BF16
weight plan required an 80 GB-class device; CPU and Metal continue to use their
expanded F32 and F16 representations respectively.
Qwen3.8 multi-token prediction defaults
Native Qwen3.8 MTP is enabled by default with one draft token on CPU, Metal, and CUDA. No environment variables are required for that default. SetIZWI_QWEN38_MTP=0 and restart or reload the model to disable MTP. The
IZWI_QWEN38_MTP_DRAFT_TOKENS setting controls proposal depth rather than
enablement; its default is 1, and supported explicit values are 1 through
3.
Managed Inference-State Support
This matrix describes ownership and provider classification in the current source tree. It is narrower than general model availability: a model can be available while a particular cache provider or hardware cell remains uncertified.
The runtime fails model loading when a required ABI-v2 operation set or exact
model/capability route is incomplete. Source/build eligibility never upgrades
a route to runtime-validated or performance-certified evidence. The runtime
does not silently switch back to a model-owned cache.
State topology by model family
The load path currently publishes ABI-v2 state as follows:
“ABI-v2 physical ownership” is not a claim of complete model-quality or
hardware performance certification. The exact model revision, capability,
backend, dtype, page geometry, attention semantics, and build feature cell must
pass its release lane.
CUDA context policy
CUDA uses the native context reported by the loaded chat checkpoint. CPU and Metal retain the configured sequence limit (4,096 tokens by default). The current native CUDA ceilings are:
Optional YaRN extensions are not included: the local adapters do not implement
their scaling parameters. Qwen3.8 fits its effective context to remaining
resources without changing the 262,144-token logical maximum. Other CUDA chat
routes reserve their selected paged KV capacity before publishing a model as
ready. Explicit context requests still fail when the requested state cannot
fit; the runtime does not overcommit VRAM.
For audio generation, automatic output budgets remain bounded. Explicit CUDA
requests can use the 32,768-token LFM2.5 Audio and Fish S2 contexts; Voxtral TTS
accepts the upstream deployment ceiling of 2,048 output frames. CPU and Metal
retain their existing limits. VibeVoice ASR uses its documented 60-minute
single-pass window on CUDA by default;
IZWI_VIBEVOICE_ASR_CUDA_MAX_AUDIO_SECS
can lower that window but cannot raise it above 60 minutes. See the
LFM2.5 Audio model card,
Fish S2 config,
Voxtral deployment,
and VibeVoice ASR model card.
Qwen3 ASR accepts up to 1,200 seconds per CUDA transcription chunk, while the
forced-aligner route accepts up to 180 seconds and rejects longer inputs
explicitly. CPU and Metal retain the checkpoint preprocessor window. Neither
CUDA route silently truncates audio at the portable preprocessor limit.
Voxtral Realtime CUDA offline decoding uses the loaded model position range
(131,072 tokens in the current checkpoint) instead of the portable 1,024-frame
service limit. Its physical paged cache rotates the 8,192-token attention
window with one spare page; inputs beyond the loaded position range fail
explicitly instead of silently dropping the tail.
Architectural processing windows are not expanded. Whisper still processes
30-second encoder windows, Kokoro still uses at most 510 phonemes per model
chunk, and streaming ASR/diarization families retain their bounded working
buffers while the runtime orchestrates longer inputs.
Cache policy support
Eligible CUDA routes automatically use resident block-table metadata,
admission-grown physical backing, device/shape-keyed attention selection,
bounded sampling readback, VRAM-tiered continuous batches, and stable one-pass
decode graph buckets. Graph keys include exact K/V and metadata tensor
identities, geometry, dtype, scale, and softcap. Partitioned decode, FP8
storage, unobserved/pre-Ampere devices, and any capture/replay failure retain
eager native decode. CPU and Metal do not enter this policy.
For configuration, counters, benchmarks, and rollback, see
KV Cache Operations.
Verification Guidance
Use the following expectations when validating a host:- Apple Silicon, macOS 15+: build or install a Metal-capable binary and run with
--backend metalorIZWI_BACKEND=metal. - macOS 12-14: run with
--backend cpu;autoand explicitmetalrequests fall back to CPU. - Linux/Windows GitHub Release: run
izwi serve --backend cpu, thenizwi status --detailed. - Docker CUDA on NVIDIA Linux hosts: run
docker compose --profile cuda up, then confirm the container selects CUDA through/v1/healthorizwi status --detailedfrom a matching client environment. Eligible Candle FlashAttention, Qwen3/Qwen3.5 RoPE, and Gemma RMSNorm CUDA routes activate automatically; their environment variables remain explicit rollback switches. Qwen3.8 is an independently gated family and must report its own execution path and evidence. - Linux/Windows source build for CUDA: build with
cargo build --release --features cuda, then run with--backend cudaorIZWI_BACKEND=cuda. Thecudawrapper includes Candle FlashAttention, whilecudnnadditionally enables matching Candle/cuDNN convolution paths.