• Text to Speech
  • Voice Cloning
  • Multilingual
  • Long-form Speech
  • Pronunciation Control

MOSS-TTS

MOSS-TTS v1.5 is a flagship speech generation model for long-form and multilingual content. It supports 31 languages, zero-shot voice cloning, and audio continuation. Fine-grained control over pronunciation, duration, and pauses brings natural, coherent, and consistent speech to narration, dubbing, and audio content.

Authors
OpenMOSS Team
Version
v1.5
Open-source License
Apache 2.0

We are introducing our next-generation flagship text-to-speech (TTS) model. It is built for real products, beyond polished demos: ready to deploy and scale, and designed to make people and businesses more productive. As a true foundation model for speech, it sets a leading standard for commercial use, open-source accessibility, and research value.

We designed the model around three essential ingredients for speech generation: a high-quality discrete audio tokenizer, large-scale pretraining data that is both high quality and diverse, and an efficient method for modeling discrete tokens. Together, they deliver state-of-the-art generation quality while keeping the architecture as simple as possible: an autoregressive approach with a lean design and powerful results.

Model Overview

The model achieves industry-leading speech synthesis quality in both objective and subjective evaluations. Built around voice cloning, it supports stable generation of very long speech, token-level duration control, multilingual and code-switching synthesis, and fine-grained pronunciation control at the pinyin and phoneme levels. These capabilities make it a production-ready foundation for scalable speech applications.

At its core, the model advances speech generation through three key factors: a high-quality audio tokenizer, large-scale, high-quality and diverse pretraining data, and an efficient method for modeling discrete tokens. Together, these ingredients enable industry-leading performance with a surprisingly simple recipe: an autoregressive approach that keeps the architecture lean while delivering powerful results.

The architecture of the MOSS Audio Tokenizer is shown below:

MOSS Audio Tokenizer

MOSS Audio Tokenizer Architecture

MOSS Audio Tokenizer is a 1.6B-parameter audio tokenizer based on the Cat (Causal Audio Tokenizer with Transformer) architecture. Designed as a unified discrete audio interface for autoregressive audio foundation models, it combines lossless reconstruction with excellent audio–text semantic alignment. For details, see MOSS Audio Tokenizer.

The two architectures used in MOSS-TTS are illustrated below.

Delay Pattern Vs. Local Transformer

MOSS-TTS Architecture - Scalable Cloud & Efficient Edge Deployment

To support both real-world deployment and academic research, MOSS-TTS trains and open-sources two complementary architectures. Left: Delay-Pattern (MossTTSDelay). Right: Global Latent + Local Transformer (MossTTSLocal). For details, see MOSS-TTS.

Benchmark

LibriSpeech Speech Evaluation Metrics (MOSS Audio Tokenizer vs. Open-source Tokenizers)

The figure below compares our MOSS Audio Tokenizer with other open-source speech tokenizers on the LibriSpeech test-clean subset. The metrics are SIM, STOI, PESQ-NB, and PESQ-WB; higher values are better. At inference time, we control the bitrate (bps) of the same model by varying the number of RVQ codebooks used.

LibriSpeech SIM comparisonLibriSpeech PESQ-NB comparisonLibriSpeech STOI comparisonLibriSpeech PESQ-WB comparison

TTS Model Comparison (Zero-shot Voice Cloning)

modelparamsopen-sourceenzh
WER/%↓SIM/%↑CER/%↓SIM/%↑
DiTAR0.6BClosed Source1.6973.51.0275.3
FishAudio-S14BClosed Source1.7262.571.2272.1
CosyVoice31.5BClosed Source2.22721.1278.1
Seed-TTS-Closed Source2.2576.21.1279.6
MiniMax-Speech-Closed Source1.6569.20.8378.3
CosyVoice0.3BOpen Source4.2960.93.6372.3
CosyVoice20.5BOpen Source3.0965.91.3875.7
CosyVoice30.5BOpen Source2.0271.81.1678
F5-TTS0.3BOpen Source2671.5376
SparkTTS0.5BOpen Source3.1457.31.5466
FireRedTTS0.5BOpen Source3.82461.5163.5
FireRedTTS-21.5BOpen Source1.9566.51.1473.6
Qwen2.5-Omni7BOpen Source2.7263.21.775.2
FishAudio-S1-mini0.5BOpen Source1.94551.1868.5
IndexTTS21.5BOpen Source2.2370.61.0376.5
VibeVoice1.5BOpen Source3.0468.91.1674.4
HiggsAudio-v23BOpen Source2.4467.71.574
GLM-TTS1.5BOpen Source2.2367.21.0376.1
GLM-TTS-RL1.5BOpen Source1.9168.10.8976.4
VoxCPM0.5BOpen Source1.8572.90.9377.2
Qwen3-TTS0.6BOpen Source1.6870.391.2376.4
Qwen3-TTS1.7BOpen Source1.571.451.3376.72
MossTTSDelay8BOpen Source1.8470.861.3776.98
MossTTSLocal1.7BOpen Source1.9373.281.4479.62

Model Demos