• Streaming generation
  • Real time
  • Voice agents

MOSS-TTS-Realtime

Context-aware, multi-turn streaming TTS: a speech generation foundation model designed for real-time voice agent interactions. Users interact through speech, and the system generates continuous, natural spoken responses in real time.

Date
2026.02.10
Author
OpenMOSS Team
Organization
MOSI.AI

Introducing MOSS-TTS-Realtime, a context-aware, multi-turn streaming text-to-speech (TTS) foundation model designed for spoken interactions. Users speak to the system, which generates continuous, natural responses in real time. Traditional TTS systems often generate each assistant reply in isolation, ignoring previous turns and the user's spoken input. This loses context and disrupts prosody. Our approach instead models the full multi-turn conversation while generating assistant speech, conditioning on both the text and acoustic representations of the user's earlier utterances. By tightly combining contextual awareness with streaming synthesis, the system produces low-latency, incremental audio responses while maintaining voice consistency and conversational coherence, resulting in highly natural, human-like speech.

Model Overview

MOSS-TTS-Realtime is a low-latency speech synthesis system built on a scalable two-stage architecture comprising a 1.7B-parameter backbone and a 200M-parameter local Transformer. The backbone models rich linguistic and contextual representations and exposes intermediate hidden states to the local Transformer, which autoregressively generates RVQ-based audio tokens. MOSS-Audio-Tokenizer then efficiently reconstructs high fidelity waveforms.

The model is trained on large-scale datasets containing more than 2.5 million hours of single-speaker speech and more than 1 million hours of two-speaker and multi-speaker conversations. This substantially improves voice consistency and coherence across turns. MOSS-TTS-Realtime supports more than ten languages in addition to Chinese and English, demonstrating strong multilingual generalization. The model supports a maximum context length of 32K tokens.

MOSS-TTS-Realtime model architecture

MOSS-TTS-Realtime model architecture

Model Demos