• Dialogue generation
  • Multiple speakers
  • Long form audio
  • Voice cloning
  • Multilingual

MOSS-TTSD

Generate complete multi-speaker conversations from character scripts, with up to five speakers, 20 languages, and 60 minutes of audio in a single generation. Set each character's voice with a few seconds of reference audio and create natural turn-taking, overlapping speech, and expressive delivery for podcasts, audiobooks, dubbing, and sports commentary.

Author
OpenMOSS Team
Version
v1.0
Open source license
Apache 2.0

MOSS-TTSD is a dialogue speech generation model for long form audio. It synthesizes highly natural and expressive multi-speaker conversations across languages. The model supports continuous long form generation and flexible control over dialogue between multiple speakers. It also offers leading zero-shot voice cloning: a short reference clip is enough to reproduce a voice with high fidelity. Designed for real content production, MOSS-TTSD can be used for podcasts, audiobooks, sports and esports commentary, dubbing, crosstalk comedy, and other entertainment formats.

The model shares the same architecture as MOSS-TTS.

View architecture

Model Overview

MOSS-TTSD (Text-To-Spoken Dialogue) is an open source model for long form dialogue speech synthesis. It generates expressive multilingual dialogue in languages including Chinese, English, Japanese, Korean, Spanish, Portuguese, French, German, Italian, Russian, and Arabic. It turns dialogue scripts into natural, human-like speech, handles multiple speakers with accurate speaker switching, and delivers rich interactions. Enhanced long context modeling enables stable, coherent generation of long audio in a single pass, making the model suitable for podcast-length productions and extended narratives. MOSS-TTSD also provides powerful zero-shot voice cloning: with only a short reference clip, it maintains a consistent speaker identity and natural conversational prosody. Applications include AI-generated podcasts, audiobook narration, sports and esports commentary, film and television dubbing, variety shows, animated character dialogue, crosstalk comedy, and other long form dialogue content.

ZH

vs doubao-podcast

Win: 35.8%Tie: 32.6%Lose: 31.6%

vs Eleven V3

Win: 44.4%Tie: 20.2%Lose: 35.4%

EN

vs gemini-2.5-flash-preview-tts

Win: 44%Tie: 22%Lose: 34%

vs gemini-2.5-pro-preview-tts

Win: 45.5%Tie: 21.2%Lose: 33.3%

vs Eleven V3

Win: 33.3%Tie: 25.3%Lose: 41.4%

WinTieLose

Model Demos