Voice · qwen/qwen3-tts
Qwen3-TTS — preset, clone and design in one model
Qwen3-TTS voices text in ten languages and combines three practical modes: preset speakers, timbre cloning from short audio, and a brand-new voice created from a plain-language description. A separate instruction controls emotion, pace and delivery.
from 3,1 ₽ per voice-over up to 2,000 chars
What Qwen3-TTS can do
- Nine preset speakers, including Serena, Aiden, Vivian, Eric and Sohee.
- Voice cloning from a clean recording as short as three seconds.
- New voice design from a description of timbre, age, accent and character.
- English, Russian and eight more languages, plus automatic detection.
Three modes for three jobs
Switch modes in the voice-over form and creator-ai reveals only the inputs that the selected mode needs.
| Mode | Input | Best for |
|---|---|---|
| Preset voice | Text, language and one of 9 speakers | Fast narration for videos, decks and announcements |
| Voice clone | Short audio; a transcript is recommended | Keeping a familiar timbre across a content series |
| Voice design | A description of timbre, age, accent and delivery | A distinctive narrator for a brand or character |
How to clone a voice
-
Prepare a sample
Record at least three seconds of clean speech with no music, echo or second speaker.
-
Add a transcript
Enter what was said. It is optional, but helps the model separate the voice from the words.
-
Enter the new script
Pick the output language and optionally add an emotion or pacing instruction.
Languages and delivery control
Qwen3-TTS supports Chinese, English, Japanese, Korean, French, German, Italian, Spanish, Portuguese and Russian. Auto mode detects the language from your text.
The style field accepts a normal phrase such as “slow and calm”, “bold and playful”, or “professional documentary narrator”. It is sent separately and does not alter the script.
Qwen3-TTS FAQ
How much reference audio do I need?
Around three seconds is enough, although a longer clean clip usually captures timbre better. Avoid background music, noise and reverb.
Can I clone in one language and speak another?
Yes. Qwen3-TTS supports cross-language cloning: the sample supplies the timbre while the target text controls the output language.
How is voice design different from a preset speaker?
A preset is selected from a fixed list. Voice design creates a new speaker from a description such as “a deep male voice with a slight British accent”.
Which audio formats can I upload?
WAV, MP3, M4A, AAC, OGG and WEBM up to 10 MB. An unprocessed, music-free recording gives the best clone.
Can I control pace and emotion?
Yes. Describe the delivery in the style field; the model understands pace, mood, volume and speaking character.