Create custom voices from text descriptions — no audio samples required.

Overview

Voice design generates unique voices based on natural language descriptions. Describe the voice you want, and Izwi creates it:
  • No samples needed — Create voices from scratch
  • Infinite variety — Design any voice you can describe
  • Quick iteration — Rapidly test different voice concepts
  • Creative freedom — Perfect for characters and personas

Getting Started

Download a Voice Design Model

Design a Voice

Describe the voice you want:

Using the Web UI

Voice design now lives inside the unified Voices workspace.

Step 1: Describe Your Voice

  1. Navigate to Voices in the sidebar and choose the design flow
  2. Enter a description of your desired voice
  3. Be specific about characteristics you want

Step 2: Generate Sample

  1. Enter sample text to hear the voice
  2. Click Generate
  3. Listen to the result

Step 3: Iterate

  • Adjust your description
  • Generate again
  • Repeat until satisfied

Voice Description Tips

Effective Descriptions

Include details about:

Example Descriptions

News anchor:
Children’s narrator:
AI assistant:
Audiobook narrator:

Using the CLI

Use izwi tts with a VoiceDesign model and pass the voice description with --instructions:
You can iterate quickly by changing only the prompt:

Using the API

Endpoint

Request

Example (curl)

Voice design history is also available through /v1/voice-designs. See the API Reference for the persisted route family and /v1/audio/speech streaming details.

Available Models

Larger models better interpret complex descriptions.

Best Practices

Be Specific

❌ “A nice voice” ✅ “A warm, professional female voice in her 40s with a calm, reassuring tone”

Use Comparisons

“Similar to a podcast host — conversational but polished”

Describe the Context

“A voice suitable for meditation apps — slow, soothing, and peaceful”

Iterate

Start broad, then refine:
  1. “A male voice”
  2. “A young male voice with energy”
  3. “A young male voice with energy, like a sports commentator”

Limitations

  • Consistency — Same description may produce slightly different voices
  • Extreme requests — Very unusual voices may not generate well
  • Accents — Some accents are better supported than others
  • Singing — Designed for speech, not singing

Voice Design vs Voice Cloning


See Also