• Lightweight Model
  • CPU Inference
  • Streaming Output
  • Voice Cloning
  • Multilingual

MOSS-TTS-Nano

MOSS-TTS-Nano delivers streaming speech generation on CPUs with roughly 100M parameters, along with voice cloning from short recordings, cross-lingual synthesis, and 48 kHz stereo input and output. It automatically chunks long text, requires neither a GPU nor additional fine-tuning for new voices, and is suited to local text-to-speech and lightweight voice applications.

Author
OpenMOSS Team
Model Size
100M Parameters
Open-Source License
Apache 2.0
MOSS-TTS-Nano — Open Source 100M TTS Model

MOSS-TTS-Nano is a deployment-first TTS model designed for real-time speech generation, voice cloning, and lightweight integration. Built on an audio tokenizer plus an autoregressive LLM pipeline, it is compact enough to run on a CPU while supporting Chinese, English, and a broad range of other languages.

The model pairs the approximately 20-million-parameter MOSS-Audio-Tokenizer-Nano with a small LLM for autoregressive token prediction. The tokenizer uses a CNN-free causal Transformer architecture and produces a 12.5 Hz token stream with 16 RVQ codebooks. It supports variable bitrates from 0.125 to 2 kbps while maintaining 48 kHz stereo output quality. Voice cloning is driven entirely by a short reference recording, with no additional fine-tuning.

Core Features

Model Size

0.1 B

Compact enough for practical CPU inference; no GPU required.

Audio Quality

48 kHz Stereo

Native stereo input and output at the full 48 kHz sample rate.

Languages

20

Chinese, English, Japanese, Korean, Spanish, French, and more.

Tiny Tokenizer

~20M Parameters

CNN-free causal Transformer, 16 RVQ codebooks, 12.5 Hz.

Bitrate

0.125–2 kbps

Variable bitrate through a configurable number of codebooks.

Inference

Real-Time

Streaming inference with low time to first token and automatic chunking for long text.

Model Architecture

Following the family-level design described in the MOSS-TTS technical report, the model combines discrete audio tokens, autoregressive modeling, and large-scale pretraining. The Nano variant retains this approach in a smaller, deployment-first configuration. MOSS-TTS technical report

MOSS-Audio-Tokenizer-Nano

MOSS-Audio-Tokenizer-Nano Architecture

The tokenizer is a causal Transformer audio codec that compresses 48 kHz stereo audio into a 12.5 fps RVQ token stream for scalable autoregressive modeling. Its encoder and decoder each contain 12 causal Transformer blocks with sliding-window attention. The quantizer uses 16 RVQ layers, keeping token sequences compact enough for long-context generation.

20M params · 48kHz stereo · 16-layer RVQ

MOSS-TTS-Nano

MOSS-TTS-Nano Architecture

On top of the tokenizer, MOSS-TTS-Nano uses hierarchical token modeling based on a Local Transformer. Instead of RVQ-aware time delays, it sums the embeddings of all RVQ layers at each aligned time step and feeds the hidden state into a single Transformer backbone. The backbone produces a global latent variable at each step, which a lightweight autoregressive Local Transformer expands into an intra-step token block, predicting one text/filler token followed by 16 RVQ audio tokens.

100 M params · Local Transformer · Tiny, Fast and Powerful

Model Demos

Language Coverage

Japanese, Korean, Spanish, French, German, Italian, Hungarian, Russian, Persian, Arabic, Polish, Portuguese, Czech, Danish, Swedish, Greek, and Turkish, in addition to Chinese and English.