Logo

Voice · qwen/qwen3-tts

Qwen3-TTS — preset, clone and design in one model

Qwen3-TTS voices text in ten languages and combines three practical modes: preset speakers, timbre cloning from short audio, and a brand-new voice created from a plain-language description. A separate instruction controls emotion, pace and delivery.

from 3,1 ₽ per voice-over up to 2,000 chars

What Qwen3-TTS can do

  • Nine preset speakers, including Serena, Aiden, Vivian, Eric and Sohee.
  • Voice cloning from a clean recording as short as three seconds.
  • New voice design from a description of timbre, age, accent and character.
  • English, Russian and eight more languages, plus automatic detection.

Three modes for three jobs

Switch modes in the voice-over form and creator-ai reveals only the inputs that the selected mode needs.

Mode Input Best for
Preset voice Text, language and one of 9 speakers Fast narration for videos, decks and announcements
Voice clone Short audio; a transcript is recommended Keeping a familiar timbre across a content series
Voice design A description of timbre, age, accent and delivery A distinctive narrator for a brand or character

How to clone a voice

  1. Prepare a sample

    Record at least three seconds of clean speech with no music, echo or second speaker.

  2. Add a transcript

    Enter what was said. It is optional, but helps the model separate the voice from the words.

  3. Enter the new script

    Pick the output language and optionally add an emotion or pacing instruction.

Languages and delivery control

Qwen3-TTS supports Chinese, English, Japanese, Korean, French, German, Italian, Spanish, Portuguese and Russian. Auto mode detects the language from your text.

The style field accepts a normal phrase such as “slow and calm”, “bold and playful”, or “professional documentary narrator”. It is sent separately and does not alter the script.

Qwen3-TTS FAQ

How much reference audio do I need?

Around three seconds is enough, although a longer clean clip usually captures timbre better. Avoid background music, noise and reverb.

Can I clone in one language and speak another?

Yes. Qwen3-TTS supports cross-language cloning: the sample supplies the timbre while the target text controls the output language.

How is voice design different from a preset speaker?

A preset is selected from a fixed list. Voice design creates a new speaker from a description such as “a deep male voice with a slight British accent”.

Which audio formats can I upload?

WAV, MP3, M4A, AAC, OGG and WEBM up to 10 MB. An unprocessed, music-free recording gives the best clone.

Can I control pace and emotion?

Yes. Describe the delivery in the style field; the model understands pace, mood, volume and speaking character.