| Voice | TTS + ASR + Chat model (or unified LFM2.5-Audio-1.5B-GGUF) |
| Chat | Chat model (Qwen3, Qwen3.5, LFM2.5, or Gemma) |
| Speaker Attributed ASR | Granite-Speech-4.1-2B-Plus |
| Voices | Built-in voice model for presets; Base or VibeVoice model for cloning; VoiceDesign model for design |
| Text-to-Speech | TTS model |
| Studio | TTS model |
| Settings | No model required |
| Transcription | ASR model (Parakeet-TDT-0.6B-v3 default; Qwen3/Whisper/Granite Speech/LFM2.5 also supported) |
| Diarization | diar_streaming_sortformer_4spk-v2.1 (+ optional ASR and aligner models) |
| Forced Alignment | Qwen3-ForcedAligner-0.6B (or -4bit) |
| Voice Cloning | Qwen3 TTS Base model (Qwen3-TTS-12Hz-*-Base*) |
| Voice Design | Qwen3 TTS VoiceDesign model (Qwen3-TTS-12Hz-1.7B-VoiceDesign*) |