We are introducing our next-generation flagship text-to-speech (TTS) model. It is built for real products, beyond polished demos: ready to deploy and scale, and designed to make people and businesses more productive. As a true foundation model for speech, it sets a leading standard for commercial use, open-source accessibility, and research value.
We designed the model around three essential ingredients for speech generation: a high-quality discrete audio tokenizer, large-scale pretraining data that is both high quality and diverse, and an efficient method for modeling discrete tokens. Together, they deliver state-of-the-art generation quality while keeping the architecture as simple as possible: an autoregressive approach with a lean design and powerful results.
Model Overview
The model achieves industry-leading speech synthesis quality in both objective and subjective evaluations. Built around voice cloning, it supports stable generation of very long speech, token-level duration control, multilingual and code-switching synthesis, and fine-grained pronunciation control at the pinyin and phoneme levels. These capabilities make it a production-ready foundation for scalable speech applications.
At its core, the model advances speech generation through three key factors: a high-quality audio tokenizer, large-scale, high-quality and diverse pretraining data, and an efficient method for modeling discrete tokens. Together, these ingredients enable industry-leading performance with a surprisingly simple recipe: an autoregressive approach that keeps the architecture lean while delivering powerful results.
The architecture of the MOSS Audio Tokenizer is shown below:
MOSS Audio Tokenizer

MOSS Audio Tokenizer is a 1.6B-parameter audio tokenizer based on the Cat (Causal Audio Tokenizer with Transformer) architecture. Designed as a unified discrete audio interface for autoregressive audio foundation models, it combines lossless reconstruction with excellent audio–text semantic alignment. For details, see MOSS Audio Tokenizer.
The two architectures used in MOSS-TTS are illustrated below.
Delay Pattern Vs. Local Transformer

To support both real-world deployment and academic research, MOSS-TTS trains and open-sources two complementary architectures. Left: Delay-Pattern (MossTTSDelay). Right: Global Latent + Local Transformer (MossTTSLocal). For details, see MOSS-TTS.
Benchmark
LibriSpeech Speech Evaluation Metrics (MOSS Audio Tokenizer vs. Open-source Tokenizers)
The figure below compares our MOSS Audio Tokenizer with other open-source speech tokenizers on the LibriSpeech test-clean subset. The metrics are SIM, STOI, PESQ-NB, and PESQ-WB; higher values are better. At inference time, we control the bitrate (bps) of the same model by varying the number of RVQ codebooks used.




TTS Model Comparison (Zero-shot Voice Cloning)
| model | params | open-source | en | zh | ||
|---|---|---|---|---|---|---|
| WER/%↓ | SIM/%↑ | CER/%↓ | SIM/%↑ | |||
| DiTAR | 0.6B | Closed Source | 1.69 | 73.5 | 1.02 | 75.3 |
| FishAudio-S1 | 4B | Closed Source | 1.72 | 62.57 | 1.22 | 72.1 |
| CosyVoice3 | 1.5B | Closed Source | 2.22 | 72 | 1.12 | 78.1 |
| Seed-TTS | - | Closed Source | 2.25 | 76.2 | 1.12 | 79.6 |
| MiniMax-Speech | - | Closed Source | 1.65 | 69.2 | 0.83 | 78.3 |
| CosyVoice | 0.3B | Open Source | 4.29 | 60.9 | 3.63 | 72.3 |
| CosyVoice2 | 0.5B | Open Source | 3.09 | 65.9 | 1.38 | 75.7 |
| CosyVoice3 | 0.5B | Open Source | 2.02 | 71.8 | 1.16 | 78 |
| F5-TTS | 0.3B | Open Source | 2 | 67 | 1.53 | 76 |
| SparkTTS | 0.5B | Open Source | 3.14 | 57.3 | 1.54 | 66 |
| FireRedTTS | 0.5B | Open Source | 3.82 | 46 | 1.51 | 63.5 |
| FireRedTTS-2 | 1.5B | Open Source | 1.95 | 66.5 | 1.14 | 73.6 |
| Qwen2.5-Omni | 7B | Open Source | 2.72 | 63.2 | 1.7 | 75.2 |
| FishAudio-S1-mini | 0.5B | Open Source | 1.94 | 55 | 1.18 | 68.5 |
| IndexTTS2 | 1.5B | Open Source | 2.23 | 70.6 | 1.03 | 76.5 |
| VibeVoice | 1.5B | Open Source | 3.04 | 68.9 | 1.16 | 74.4 |
| HiggsAudio-v2 | 3B | Open Source | 2.44 | 67.7 | 1.5 | 74 |
| GLM-TTS | 1.5B | Open Source | 2.23 | 67.2 | 1.03 | 76.1 |
| GLM-TTS-RL | 1.5B | Open Source | 1.91 | 68.1 | 0.89 | 76.4 |
| VoxCPM | 0.5B | Open Source | 1.85 | 72.9 | 0.93 | 77.2 |
| Qwen3-TTS | 0.6B | Open Source | 1.68 | 70.39 | 1.23 | 76.4 |
| Qwen3-TTS | 1.7B | Open Source | 1.5 | 71.45 | 1.33 | 76.72 |
| MossTTSDelay | 8B | Open Source | 1.84 | 70.86 | 1.37 | 76.98 |
| MossTTSLocal | 1.7B | Open Source | 1.93 | 73.28 | 1.44 | 79.62 |