
Understanding audio involves much more than transcribing words. It also requires perceiving acoustic cues, identifying speakers and emotions, interpreting environmental sounds, reasoning about temporal context, and handling complex multi-step inferences. MOSS-Audio aims to bring these capabilities together in one model.
This release includes four models: MOSS-Audio-4B-Instruct, MOSS-Audio-4B-Thinking, MOSS-Audio-8B-Instruct, and MOSS-Audio-8B-Thinking. The Instruct variants are optimized for instruction following, while the Thinking variants offer stronger chain-of-thought reasoning.
Core Capabilities
Speech Recognition
ASR + Timestamps
Transcribes speech across a range of acoustic conditions, with optional word-level and sentence-level timestamp alignment.
Speaker Analysis
Identity and Emotion
Identifies speaker characteristics, analyzes emotional states, and detects key acoustic events.
Environmental Audio
Scene Understanding
Extracts cues from background sounds, noise, and non-speech signals to infer scene context.
Music Understanding
Style and Emotion
Analyzes musical style, emotional progression, instrumentation, and salient acoustic features.
Audio Question Answering
Open-Ended Q&A
Answers questions about speech, podcasts, meetings, and recordings, and generates summaries.
Complex Reasoning
Chain of Thought
Performs multi-step reasoning about audio content through chain-of-thought training and reinforcement learning.
Model Architecture
MOSS-Audio has a modular design with three components: a dedicated audio encoder, a modality adapter, and a large language model. Raw audio is encoded into a continuous temporal representation at 12.5 Hz and projected into the LLM embedding space for autoregressive text generation.
Overall architecture of MOSS-Audio
DeepStack Cross-Layer Feature Injection
Using only the encoder’s top-layer features loses low-level prosody, transient events, and local time-frequency structure. MOSS-Audio uses a DeepStack cross-layer injection module: features from earlier and intermediate encoder layers are projected independently and injected into the early LLM layers, preserving information at multiple levels of granularity, from acoustic details to high-level semantic abstractions.
Multi-layer injection · Preserves prosody & transients · Encoder trained from scratch
Time-Aware Representation
During pretraining, explicit temporal marker tokens are inserted between audio-frame representations at fixed intervals. This lets the model learn “what happened when” within a unified text-generation framework. It naturally supports timestamped ASR, event localization, time-based question answering, and retrieval across long audio recordings.
Time-marker insertion · 12.5 Hz token stream · Qwen3-4B backbone
Evaluation Highlights
MOSS-Audio was evaluated on a comprehensive set of audio-understanding benchmarks covering general audio, speech description, ASR, and timestamp alignment.
General Audio (Average Accuracy)
71.08
MOSS-Audio-8B-Thinking achieves an average accuracy of 71.08, outperforming all open-source models.
Speech Description (LLM Judge)
3.7252 / 5
MOSS-Audio-Instruct leads on 11 of the 13 speech-description dimensions.
ASR (Overall CER ↓)
11.30
MOSS-Audio achieves the lowest overall CER across 12 ASR evaluation dimensions.
Timestamped ASR · AAS ↓
35.77 / 131.61
MOSS-Audio-8B-Instruct achieves an AAS of 35.77 on AISHELL-1 and 131.61 on LibriSpeech.

General Audio Understanding accuracy comparison across open-source and closed-source models
Speech Description
Fine-grained speech-style descriptions are evaluated across 13 descriptive dimensions using an LLM-as-judge protocol.
LLM-Judge Score ↑

ASR
Summary of CER results across 12 ASR evaluation dimensions. Lower is better.
CER ↓
| Model | Overall | Health | Dialect | Singing | Non-Speech | Code-Switch | Clean | Noisy | Whisper | Far/Near | Multi-Spk | Age | Semantic |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Paraformer-Large | 15.77 | 22.18 | 43.45 | 32.34 | 4.95 | 12.65 | 3.11 | 4.67 | 5.02 | 17.46 | 20.33 | 14.96 | 7.14 |
| GLM-ASR-Nano | 17.29 | 24.49 | 22.39 | 51.95 | 4.65 | 11.88 | 3.68 | 5.02 | 4.94 | 27.51 | 28.02 | 17.19 | 7.32 |
| Fun-ASR-Nano | 12.04 | 21.99 | 7.80 | 19.35 | 4.76 | 11.23 | 2.98 | 3.46 | 3.78 | 18.38 | 19.82 | 14.95 | 6.08 |
| SenseVoice-Small | 14.50 | 24.04 | 8.89 | 23.79 | 4.92 | 13.90 | 4.13 | 4.93 | 5.57 | 26.66 | 24.06 | 17.63 | 7.55 |
| Kimi-Audio-7B-Instruct | 14.12 | 21.11 | 29.34 | 21.76 | 4.68 | 16.38 | 2.20 | 2.15 | 2.66 | 21.02 | 20.61 | 16.74 | 6.12 |
| Qwen2.5-Omni-3B | 15.26 | 24.65 | 33.87 | 24.24 | 5.54 | 11.66 | 2.76 | 3.56 | 4.32 | 22.15 | 22.91 | 15.17 | 7.24 |
| Qwen2.5-Omni-7B | 15.05 | 23.85 | 31.91 | 22.69 | 4.56 | 12.97 | 2.52 | 3.16 | 3.64 | 25.38 | 21.01 | 16.13 | 6.78 |
| Qwen3-Omni-30B-A3B-Instruct | 11.39 | 20.73 | 15.63 | 16.01 | 4.73 | 11.30 | 2.23 | 2.47 | 1.90 | 17.08 | 18.15 | 11.46 | 5.74 |
| MOSS-Audio-4B-Instruct | 11.58 | 21.11 | 11.84 | 10.79 | 4.01 | 10.11 | 3.11 | 3.72 | 3.29 | 18.48 | 20.33 | 15.09 | 8.15 |
| MOSS-Audio-8B-Instruct | 11.30 | 19.18 | 8.76 | 9.81 | 4.31 | 10.18 | 2.70 | 3.20 | 2.75 | 24.04 | 24.36 | 15.26 | 7.69 |
Timestamped ASR
Timestamp alignment quality is measured by AAS on Chinese and English benchmarks. Lower is better.
AAS ↓
| Model | AISHELL-1 (zh) | LibriSpeech (en) |
|---|---|---|
| Qwen3-Omni-30B-A3B-Instruct | 833.66 | 646.95 |
| Gemini-3.1-Pro | 708.24 | 871.19 |
| MOSS-Audio-4B-Instruct | 76.96 | 358.13 |
| MOSS-Audio-8B-Instruct | 35.77 | 131.61 |
Model Demos
Released Models
| Model | Audio Encoder | LLM Backbone | Total Size | Hugging Face | ModelScope |
|---|---|---|---|---|---|
| MOSS-Audio-4B-Instruct | MOSS-Audio-Encoder | Qwen3-4B | ~4.6B | Model | Model |
| MOSS-Audio-4B-Thinking | MOSS-Audio-Encoder | Qwen3-4B | ~4.6B | Model | Model |
| MOSS-Audio-8B-Instruct | MOSS-Audio-Encoder | Qwen3-8B | ~8.6B | Model | Model |
| MOSS-Audio-8B-Thinking | MOSS-Audio-Encoder | Qwen3-8B | ~8.6B | Model | Model |