Qwen3.8-LiveTranslate: Names the speaker. Carries the meaning.
The model reduces average lagging to 2.3 seconds across 60 input languages while adding real-time multi-speaker separation and voice cloning.
- Uses a Hybrid-MoE Thinker–Talker architecture that processes audio, video, source text, and translation in a single interleaved stream.
- Supports audio input and text output in 60 languages, with synthesized speech output in 29 languages.
- Performs speaker separation during multi-speaker dialogues to maintain individual vocal timbres in the translated audio.
- Draws on prior conversation turns and visual inputs for long-context disambiguation of proper nouns and terminology.
- Available via the DashScope WebSocket API under the model identifier qwen3.8-livetranslate-flash-realtime.
Developers building live interpretation tools can stream multi-speaker speech translations with voice cloning without managing separate ASR and diarization pipelines.
Sources
Read this as text
Back to the AI news