0.1.0-beta-18; it does not imply that every model/backend combination has
completed production hardware certification.
Configuration migration
Existing configuration files that omit the new cache fields continue to load. They resolve to a 64-token page, densefloat16 KV, and prefix reuse disabled:
EngineCoreConfig uses the equivalent field name block_size.
Both public configuration surfaces default to 64. Environment-based backend
selection is resolved separately and does not change serialized defaults.
float16, bfloat16, and float32 are accepted dense KV requests. int8 and
q4 still deserialize so an old file produces a useful startup error, but they
are rejected before model readiness. They are not silently stored as dense KV.
Rust API migration
EngineConfig remains available from the crate root for the beta-17
compatibility window. New integrations should import cache policy types from
their owning module and resolve the policy before announcing readiness:
ManagedKvRuntimeSnapshot for downstream cache observability. ABI-v2
plans, arenas, leases, transactions, and operation registries are internal
ownership types and are not replacements for removed proof-of-concept manager
handles. The additive runtime.kv_cache_policy health object is the supported
HTTP migration surface.
Requested and effective policy
Inspectruntime.kv_cache_policy in GET /v1/health. The response separates
what was requested from what the runtime enforces:
fallback_reason is an expected, explicit capacity adjustment—not an implicit
provider fallback. If policy resolution cannot preserve at least one
maximum-length request, startup fails.
Prefix isolation and capacity
Prefix reuse starts disabled. To enable it, set all three fields:[runtime] keys (enable_prefix_caching,
managed_prefix_cache_salt, max_prefix_cache_pages) or environment variables
(IZWI_ENABLE_PREFIX_CACHING, IZWI_MANAGED_PREFIX_CACHE_SALT,
IZWI_MAX_PREFIX_CACHE_PAGES). Note that a model must also opt in through its
cache contract: dense families (Qwen3, Gemma3, Voxtral, VibeVoice) participate,
while hybrid families (Qwen3.5, Qwen3.8) currently declare prefix reuse
disabled because their recurrent and convolution state has no checkpoint
boundary for shared spans.
Use a stable, non-secret namespace that changes whenever tenants or deployments
must not share cache entries. Never reuse one namespace across mutually
untrusted tenants. Rotate it after tokenizer, adapter, prompt-template, position,
or multimodal preprocessing changes if the model generation identity does not
already change.
Capacity is measured in physical pages. The effective prefix bound is the
smaller of max_prefix_cache_pages and the pages left after reserving one
max_sequence_length request. Watch these runtime fields:
counters.prefix_hits,prefix_misses,prefix_evictions,prefix_copy_on_write_pages,prefix_rejections, andprefix_retained_pages;totals.coordinator.allocated_pages,free_pages,prefix_refs,execution_pins, andactive_transactions;totals.physical_bytesandmemory_accounting, which should readphysical_arena_backing.
CUDA build and provider controls
The feature names differ slightly by crate:
Examples:
1/true/yes/on disables Optimized, while
0/false/no/off enables normal promotion. Invalid values fail startup.
Benchmark methodology
The checked-in microbenchmark uses the public physical arena ABI and emits JSON Lines. It measures synchronized operation latency; it is not end-to-end model throughput.unsupported records
for unavailable hardware. Run the additional executable axes explicitly until
the matrix runner covers them. Model-route parity and release soaks remain
separate gates; do not substitute this microbenchmark for them.
No repository benchmark artifact currently establishes a universal hard
latency target. Provider promotion must remain tied to reviewed results for the
exact hardware and shape cell.
Rollback runbook
Rollback is policy-only; ABI v1 and model-owned KV are not available.- Disable prefix reuse with
enable_prefix_caching: false. Restart and confirm health reportsprefix.mode: disabled. - On CUDA, set
IZWI_KV_DISABLE_OPTIMIZED_PROVIDER=1. Restart and confirm the exact route selects Portable rather than Optimized. - If the CUDA backend itself is suspect, explicitly select CPU or Metal and
confirm
requested_backend_availableandselected_backendin health. - Drain or restart the process before changing page size or dtype; never mix old physical allocations with a new policy.
- Preserve the failing request, health response, runtime KV snapshot, metrics, logs, candidate SHA, model revision, and hardware manifest.
Known unsupported or uncertified areas
- Configurable
int8andq4physical KV storage and attention kernels. - FP8 E4M3 CUDA KV promotion; the implementation is source-complete but no hardware/shape cell is certified yet, so dense KV remains authoritative.
- Cross-namespace or cross-tenant prefix sharing.
- Host-offloaded/tiered KV storage and distributed/multi-node cache ownership.
- Treating native release binaries for Linux or Windows as CUDA builds; those published artifacts remain CPU-only.
- Claiming CUDA device correctness from compile/link CI alone.
- Claiming a provider Optimized without numerical and measured performance certification for the exact route.