MOSS-TTSD is a dialogue speech generation model for long form audio. It synthesizes highly natural and expressive multi-speaker conversations across languages. The model supports continuous long form generation and flexible control over dialogue between multiple speakers. It also offers leading zero-shot voice cloning: a short reference clip is enough to reproduce a voice with high fidelity. Designed for real content production, MOSS-TTSD can be used for podcasts, audiobooks, sports and esports commentary, dubbing, crosstalk comedy, and other entertainment formats.
The model shares the same architecture as MOSS-TTS.
Model Overview
MOSS-TTSD (Text-To-Spoken Dialogue) is an open source model for long form dialogue speech synthesis. It generates expressive multilingual dialogue in languages including Chinese, English, Japanese, Korean, Spanish, Portuguese, French, German, Italian, Russian, and Arabic. It turns dialogue scripts into natural, human-like speech, handles multiple speakers with accurate speaker switching, and delivers rich interactions. Enhanced long context modeling enables stable, coherent generation of long audio in a single pass, making the model suitable for podcast-length productions and extended narratives. MOSS-TTSD also provides powerful zero-shot voice cloning: with only a short reference clip, it maintains a consistent speaker identity and natural conversational prosody. Applications include AI-generated podcasts, audiobook narration, sports and esports commentary, film and television dubbing, variety shows, animated character dialogue, crosstalk comedy, and other long form dialogue content.
ZH
Win: 35.8%Tie: 32.6%Lose: 31.6%
Win: 44.4%Tie: 20.2%Lose: 35.4%
EN
Win: 44%Tie: 22%Lose: 34%
Win: 45.5%Tie: 21.2%Lose: 33.3%
Win: 33.3%Tie: 25.3%Lose: 41.4%