Installation Issues
macOS: “Izwi can’t be opened because it is from an unidentified developer”
The app isn’t code-signed yet:- Go to System Settings → Privacy & Security
- Scroll down to find the Izwi message
- Click Open Anyway
macOS: Command not found: izwi
The CLI tools aren’t in your PATH:Linux: Permission denied installing .deb
Use sudo:Windows: SmartScreen blocks installation
- Click More info
- Click Run anyway
Server Issues
Server won’t start
Check if port is in use:Can’t connect to server
-
Verify server is running:
-
Check the operational probes:
-
Check the correct URL:
- Check firewall settings
Server crashes on startup
Check logs:- Insufficient memory
- Corrupted model files
- Missing dependencies
Model Issues
Model download fails
Network issues:- Check internet connection
- Try again (downloads resume automatically)
- Use a VPN if region-blocked
Model won’t load
Insufficient memory: Check available RAM:Qwen3.8-27B-FP8 on CUDA, inspect the loaded entry returned by
/v1/health. A successful compressed load reports
family_diagnostics.resident_representation as
q8_0_requantized_projections_with_dense_bf16 and
fp8_execution_mode as q8_0_compressed_fallback. This is intended for
40/48 GB-class devices with a resource-fitted context, but admission can still
fail when free VRAM or allocator headroom is insufficient. It is not a native
FP8 mode; see the support matrix.
For Qwen3.8 responses that stop with No finite Qwen3.8 sampling distribution,
check family_diagnostics.optimization_evidence.cuda_kv_storage in the loaded
model diagnostics. BF16 model activations must retain their exponent range:
F16 KV conversion can turn a finite value into infinity and corrupt subsequent
attention. Supported CUDA devices (observed compute capability 8.0+) now default
to storage_dtype: bf16 and selected_provider: cuda_bf16, with the same KV
memory footprint. Rebuild/restart and reload the model to apply the policy.
Remove an explicit IZWI_QWEN38_CUDA_BF16_KV=0 override to use the default;
that override remains available for controlled F16 comparisons. Sampling
failures include the target/draft/bonus stage and bounded numerical diagnostics;
retain those details if a failure recurs under BF16 KV.
If the diagnostic reports phase=draft and finite=0, the optional MTP head
has produced an unusable proposal. Izwi now discards that entire speculative
round, restores its cache position and draft RNG, and uses target-only sampling
for the rest of that request. This recovery is automatic with MTP enabled;
other requests retain their own MTP policy. The warning records the failing
position and draft depth, and
optimization_evidence.counters.mtp_nonfinite_draft_fallbacks_total counts
requests switched to target-only sampling. MTP cache maintenance continues to
preserve the loaded adapter’s state contract, so recovery does not remove all
MTP computation.
It prevents an invalid optional draft from aborting a healthy target stream;
it does not establish or repair the underlying source of the MTP NaNs.
Target/bonus numerical failures and backend execution failures remain errors.
For long-context requests, inspect runtime_metrics.kv_cache.models in the
health/admin diagnostics. single_sequence_token_capacity is the largest
sequence the fitted pools can retain, while full_context_sequence_capacity
is how many such sequences fit concurrently. Each arena also reports
token_capacity, full-request page claims, and workspace budget/high-water
bytes. Izwi reserves the exact prompt plus requested maximum output logically
before dispatch; reduce max_tokens or concurrency when that complete demand
does not fit. It does not evict arbitrary tokens from an active full-attention
sequence.
CUDA_ERROR_OUT_OF_MEMORY should not be returned for ordinary managed-capacity
pressure. If it appears after this version, capture the exact Git SHA, loaded
model diagnostics, the managed-KV snapshots before/after the request, and the
CUDA driver/device profile; treat it as an allocator/runtime defect rather than
raising the advertised context limit.
Corrupted model:
Model not detected after manual download
-
Verify correct directory:
- macOS:
~/Library/Application Support/izwi/models/ - Linux:
~/.local/share/izwi/models/ - Windows:
%APPDATA%\izwi\models\
- macOS:
- Check folder name matches expected variant name
-
Restart the server:
Performance Issues
Inference is slow
Use GPU acceleration: macOS (Metal):--features cuda,flash-attn or --features cuda,cudnn.
Use smaller models:
Qwen3-TTS-12Hz-0.6B-Baseinstead ofQwen3-TTS-12Hz-1.7B-Base- Quantized variants (
-4bit)
High memory usage
Unload unused models:Audio generation stutters
- Ensure models are fully loaded before use
- Use streaming mode for long text
- Check system resources
Audio Issues
No audio output
Check system audio:- Verify speakers/headphones are connected
- Check system volume
- Test with another application
Poor transcription quality
Improve audio quality:- Use a better microphone
- Reduce background noise
- Speak clearly
Microphone not detected (Web UI)
- Check browser permissions for microphone access
- Ensure correct input device is selected in system settings
- Try a different browser
GPU Issues
Metal not working (macOS)
Verify Apple Silicon:CUDA not detected (Linux/Windows)
Check NVIDIA drivers:Web UI Issues
UI won’t load
-
Verify server is running:
-
Check the URL:
http://localhost:8080 - Clear browser cache
- Try incognito/private mode
UI shows “No models loaded”
-
Download a model:
-
Load the model:
- Refresh the page
Features not working
Ensure required models are loaded:API Issues
401 Unauthorized
Izwi doesn’t require authentication by default. If you’re getting this error:- Check you’re connecting to the right server
- Verify no proxy is interfering
404 Not Found
Check the endpoint URL:- TTS:
POST /v1/audio/speech - Transcription:
POST /v1/audio/transcriptions - Chat:
POST /v1/chat/completions
500 Internal Server Error
Check server logs:- Model not loaded
- Invalid request format
- Insufficient memory
Getting More Help
Collect diagnostic information
Check logs
Report issues
Open an issue on GitHub with:- Izwi version (
izwi version --full) - Operating system and version
- Steps to reproduce
- Error messages and logs