• Voice design
  • Expressive speech
  • Emotion control
  • Chinese and English
  • Zero-shot generation

MOSS-VoiceGenerator

An open source voice design model that turns natural language descriptions of perceived age, vocal texture, accent, and emotion into a voice that speaks your chosen lines. It needs no reference recording, supports speech in Chinese and English, and offers fine-grained control over delivery. Use it for standalone speech creation or to provide a voice reference for downstream TTS.

Author
OpenMOSS Team
Open source license
Apache 2.0

Introducing MOSS-VoiceGenerator, an open source voice generation system that creates a speaker's voice directly from a free-form text description. Users can design distinctive voices for particular characters, personalities, and emotions. The system delivers exceptionally natural emotional expression, producing speech that sounds convincingly human. It supports a wide range of uses, including audiobooks, game voice acting, role-playing agents, and conversational assistants. It can also serve as a voice design layer for other TTS systems, removing the need to source reference recordings and making integration easier and more effective.

The model shares the same architecture as MOSS-TTS.

View architecture

Model Overview

MOSS-VoiceGenerator uses a unified multimodal language model architecture. It combines the voice description instruction with the text to be synthesized, encodes them together, and uses the result to drive speech generation. This jointly models voice design, style control, and content synthesis. Through instruction-to-voice alignment, the model learns the relationship between text descriptions and acoustic features. It can therefore produce high fidelity speech with the requested voice, emotion, and style directly from a free-form text prompt, without reference audio.

Model Capabilities

MOSS-VoiceGenerator shows a clear advantage in subjective evaluations. Using an internal test set of 160 samples spanning a variety of voice styles, we measured three independent dimensions: (1) Overall Preference — which voice would you choose as a user? (2) Instruction Following — which generated clip better follows the instructions? These may specify gender, age, vocal texture, emotion, accent, speaking rate, and other details. (3) Naturalness — which generated clip sounds closest to real human speech? MOSS-VoiceGenerator outperformed every TTS system in the comparison that supports voices without predefined presets and custom preview text across all three dimensions.

WinTieLose

Overall Performance

vs Qwen3-TTS-vd

Win: 61.9%Tie: 5%Lose: 33.1%

vs MiniMax Voice Design

Win: 61.9%Tie: 6.9%Lose: 31.2%

vs MiMo-Audio-7B-Instruct

Win: 63.1%Tie: 10%Lose: 26.9%

Instruction Following

vs Qwen3-TTS-vd

Win: 53.1%Tie: 11.2%Lose: 35.6%

vs MiniMax Voice Design

Win: 55%Tie: 10.6%Lose: 34.4%

vs MiMo-Audio-7B-Instruct

Win: 60%Tie: 10%Lose: 30%

Naturalness

vs Qwen3-TTS-vd

Win: 57.5%Tie: 7.5%Lose: 35%

vs MiniMax Voice Design

Win: 50.3%Tie: 13.2%Lose: 36.5%

vs MiMo-Audio-7B-Instruct

Win: 59.4%Tie: 7.5%Lose: 33.1%

Model Demos