Overview
Izwi chat provides:- Local inference — Model execution stays on-device
- Multiple model families — Qwen3, Qwen3.5, Qwen3.8, LFM2.5, and Gemma
- System prompts — Shape assistant behavior
- Streaming output — Incremental response tokens
- Multimodal support (Qwen3.5 only) — Image inputs in chat API requests
Getting Started
Download a Chat Model
Start Chatting
Using the CLI
qwen3-0.6b-4bit remains the CLI default for backward compatibility.
For new setups, prefer an enabled model from izwi list, such as Qwen3-8B-GGUF or Qwen3.5-4B.
Examples:
Using the Web UI
- Open Chat in the sidebar
- Enter a prompt
- Send and review streamed output
- Switch loaded models from the model selector
Using the API
Text Chat Endpoint
Text Request Example
cURL Example
Multimodal (Image) Example
Image inputs are supported only on Qwen3.5 GGUF chat variants:stream_options.include_usage, tool-call
payloads, and strict/relaxed OpenAI compatibility profiles. See the
API Reference for the full request contract and
streaming sequence.
Supported Chat Models
Qwen3.8 is registered as its own
Qwen38Chat runtime family. It does not
inherit Qwen3.5 backend certification, state-conformance evidence, or media
capabilities. Its CUDA deployment uses a
scale-aware Q8_0 compressed fallback,
not native FP8 execution.
Multimodal Limits
- Multimodal media chat is currently limited to Qwen3.5 GGUF models.
- Video inputs are not yet implemented.
- Non-Qwen3.5 chat variants, including
Qwen3.8-27B-FP8, currently support text-only requests.
Tips
- Use
izwi listto pick a currently enabled model ID. - Use stronger models (
Qwen3-8B-GGUF,Qwen3.5-9B,Qwen3.8-27B-FP8) for harder tasks when the host has sufficient memory; check the Qwen3.8 CUDA residency requirements before deploying the 27B checkpoint. - Use smaller models (
Qwen3.5-0.8B,LFM2.5-1.2B-*) for low-latency usage.
See Also
CUDA concurrent chat
Qwen3.8 batches independent CUDA chat requests through one loaded model by default. The incremental admission policy grows cache reservations with active sequences and can suspend and replay a request when the shared cache is under pressure. It preserves automatic output limits and already streamed text. Capacity is derived from the device’s memory and model state, rather than a GPU model name. Use your normal CUDA server command; no environment flag is needed. The policy also enables scheduler-visible chunked prefill for resumable adapters. To opt out and return to conservative admission, set this before restarting the server:1 or remove the variable to restore the default.
Other model families keep their existing admission policy until they implement the
published-sequence replay contract.
Use separate conversations, or independent /v1/chat/completions requests, for
concurrent answers. Sends to the same conversation remain ordered. The playground
allows one active stream per mounted view; separate tabs can use separate conversations.
Inspect /v1/health for the effective runtime.chat_concurrency_policy, and
/v1/metrics for actual model batch widths, cache claims, suspensions and replay
work. Open HTTP streams alone do not establish concurrent model execution. Heavy
cache pressure can still queue or suspend requests; simultaneous maximum-length
histories must fit the available state budget.
Temporary batch workspace pressure is also handled before model execution: the
scheduler reduces batch width or resumable prefill size and retries the same
request without repeating streamed output. Under persistent pressure it can
suspend eligible Qwen3.8 requests. Progress from competing work keeps waiting
requests eligible; repeated failures with no progress still return an explicit
capacity error. Requests that cannot fit alone cannot be guaranteed to finish.
Capacity diagnostics group pending reservations by owner class to distinguish
model, request, and workspace promises from materialized memory.
For a controlled hardware acceptance run, use
scripts/bench/run-cuda-chat-concurrency.py as documented in
scripts/bench/README.md. It checks uncapped requests through both API routes,
staggered arrivals, cancellation, overlapping output, actual multirow forwards,
and resource release. CPU tests do not certify CUDA throughput or numerical behavior.