Introducing MOSS-VoiceGenerator, an open source voice generation system that creates a speaker's voice directly from a free-form text description. Users can design distinctive voices for particular characters, personalities, and emotions. The system delivers exceptionally natural emotional expression, producing speech that sounds convincingly human. It supports a wide range of uses, including audiobooks, game voice acting, role-playing agents, and conversational assistants. It can also serve as a voice design layer for other TTS systems, removing the need to source reference recordings and making integration easier and more effective.
The model shares the same architecture as MOSS-TTS.
Model Overview
MOSS-VoiceGenerator uses a unified multimodal language model architecture. It combines the voice description instruction with the text to be synthesized, encodes them together, and uses the result to drive speech generation. This jointly models voice design, style control, and content synthesis. Through instruction-to-voice alignment, the model learns the relationship between text descriptions and acoustic features. It can therefore produce high fidelity speech with the requested voice, emotion, and style directly from a free-form text prompt, without reference audio.
Model Capabilities
MOSS-VoiceGenerator shows a clear advantage in subjective evaluations. Using an internal test set of 160 samples spanning a variety of voice styles, we measured three independent dimensions: (1) Overall Preference — which voice would you choose as a user? (2) Instruction Following — which generated clip better follows the instructions? These may specify gender, age, vocal texture, emotion, accent, speaking rate, and other details. (3) Naturalness — which generated clip sounds closest to real human speech? MOSS-VoiceGenerator outperformed every TTS system in the comparison that supports voices without predefined presets and custom preview text across all three dimensions.
Overall Performance
Win: 61.9%Tie: 5%Lose: 33.1%
Win: 61.9%Tie: 6.9%Lose: 31.2%
Win: 63.1%Tie: 10%Lose: 26.9%
Instruction Following
Win: 53.1%Tie: 11.2%Lose: 35.6%
Win: 55%Tie: 10.6%Lose: 34.4%
Win: 60%Tie: 10%Lose: 30%
Naturalness
Win: 57.5%Tie: 7.5%Lose: 35%
Win: 50.3%Tie: 13.2%Lose: 36.5%
Win: 59.4%Tie: 7.5%Lose: 33.1%