
MOSS-TTS-Nano is a deployment-first TTS model designed for real-time speech generation, voice cloning, and lightweight integration. Built on an audio tokenizer plus an autoregressive LLM pipeline, it is compact enough to run on a CPU while supporting Chinese, English, and a broad range of other languages.
The model pairs the approximately 20-million-parameter MOSS-Audio-Tokenizer-Nano with a small LLM for autoregressive token prediction. The tokenizer uses a CNN-free causal Transformer architecture and produces a 12.5 Hz token stream with 16 RVQ codebooks. It supports variable bitrates from 0.125 to 2 kbps while maintaining 48 kHz stereo output quality. Voice cloning is driven entirely by a short reference recording, with no additional fine-tuning.
Core Features
Model Size
0.1 B
Compact enough for practical CPU inference; no GPU required.
Audio Quality
48 kHz Stereo
Native stereo input and output at the full 48 kHz sample rate.
Languages
20
Chinese, English, Japanese, Korean, Spanish, French, and more.
Tiny Tokenizer
~20M Parameters
CNN-free causal Transformer, 16 RVQ codebooks, 12.5 Hz.
Bitrate
0.125–2 kbps
Variable bitrate through a configurable number of codebooks.
Inference
Real-Time
Streaming inference with low time to first token and automatic chunking for long text.
Model Architecture
Following the family-level design described in the MOSS-TTS technical report, the model combines discrete audio tokens, autoregressive modeling, and large-scale pretraining. The Nano variant retains this approach in a smaller, deployment-first configuration. MOSS-TTS technical report
MOSS-Audio-Tokenizer-Nano

The tokenizer is a causal Transformer audio codec that compresses 48 kHz stereo audio into a 12.5 fps RVQ token stream for scalable autoregressive modeling. Its encoder and decoder each contain 12 causal Transformer blocks with sliding-window attention. The quantizer uses 16 RVQ layers, keeping token sequences compact enough for long-context generation.
20M params · 48kHz stereo · 16-layer RVQ
MOSS-TTS-Nano

On top of the tokenizer, MOSS-TTS-Nano uses hierarchical token modeling based on a Local Transformer. Instead of RVQ-aware time delays, it sums the embeddings of all RVQ layers at each aligned time step and feeds the hidden state into a single Transformer backbone. The backbone produces a global latent variable at each step, which a lightweight autoregressive Local Transformer expands into an intra-step token block, predicting one text/filler token followed by 16 RVQ audio tokens.
100 M params · Local Transformer · Tiny, Fast and Powerful
Model Demos
Language Coverage
Japanese, Korean, Spanish, French, German, Italian, Hungarian, Russian, Persian, Arabic, Polish, Portuguese, Czech, Danish, Swedish, Greek, and Turkish, in addition to Chinese and English.