• Audio Understanding
  • Speech Recognition
  • Chain-of-Thought Reasoning

MOSS-Audio

An open-source audio understanding model supporting speech recognition, environmental sound analysis, music understanding, time-aware question answering, and complex multi-step reasoning.

Date
2026.04.14
Author
OpenMOSS Team
Organization
MOSI.AI
MOSS-Audio — Open Source Audio Understanding Model

Understanding audio involves much more than transcribing words. It also requires perceiving acoustic cues, identifying speakers and emotions, interpreting environmental sounds, reasoning about temporal context, and handling complex multi-step inferences. MOSS-Audio aims to bring these capabilities together in one model.

This release includes four models: MOSS-Audio-4B-Instruct, MOSS-Audio-4B-Thinking, MOSS-Audio-8B-Instruct, and MOSS-Audio-8B-Thinking. The Instruct variants are optimized for instruction following, while the Thinking variants offer stronger chain-of-thought reasoning.

Core Capabilities

Speech Recognition

ASR + Timestamps

Transcribes speech across a range of acoustic conditions, with optional word-level and sentence-level timestamp alignment.

Speaker Analysis

Identity and Emotion

Identifies speaker characteristics, analyzes emotional states, and detects key acoustic events.

Environmental Audio

Scene Understanding

Extracts cues from background sounds, noise, and non-speech signals to infer scene context.

Music Understanding

Style and Emotion

Analyzes musical style, emotional progression, instrumentation, and salient acoustic features.

Audio Question Answering

Open-Ended Q&A

Answers questions about speech, podcasts, meetings, and recordings, and generates summaries.

Complex Reasoning

Chain of Thought

Performs multi-step reasoning about audio content through chain-of-thought training and reinforcement learning.

Model Architecture

MOSS-Audio has a modular design with three components: a dedicated audio encoder, a modality adapter, and a large language model. Raw audio is encoded into a continuous temporal representation at 12.5 Hz and projected into the LLM embedding space for autoregressive text generation.

MOSS-Audio Architecture: Audio Encoder → Modality Adapter → LLM

Overall architecture of MOSS-Audio

DeepStack Cross-Layer Feature Injection

Using only the encoder’s top-layer features loses low-level prosody, transient events, and local time-frequency structure. MOSS-Audio uses a DeepStack cross-layer injection module: features from earlier and intermediate encoder layers are projected independently and injected into the early LLM layers, preserving information at multiple levels of granularity, from acoustic details to high-level semantic abstractions.

Multi-layer injection · Preserves prosody & transients · Encoder trained from scratch

Time-Aware Representation

During pretraining, explicit temporal marker tokens are inserted between audio-frame representations at fixed intervals. This lets the model learn “what happened when” within a unified text-generation framework. It naturally supports timestamped ASR, event localization, time-based question answering, and retrieval across long audio recordings.

Time-marker insertion · 12.5 Hz token stream · Qwen3-4B backbone

Evaluation Highlights

MOSS-Audio was evaluated on a comprehensive set of audio-understanding benchmarks covering general audio, speech description, ASR, and timestamp alignment.

General Audio (Average Accuracy)

71.08

MOSS-Audio-8B-Thinking achieves an average accuracy of 71.08, outperforming all open-source models.

Speech Description (LLM Judge)

3.7252 / 5

MOSS-Audio-Instruct leads on 11 of the 13 speech-description dimensions.

ASR (Overall CER ↓)

11.30

MOSS-Audio achieves the lowest overall CER across 12 ASR evaluation dimensions.

Timestamped ASR · AAS ↓

35.77 / 131.61

MOSS-Audio-8B-Instruct achieves an AAS of 35.77 on AISHELL-1 and 131.61 on LibriSpeech.

Model evaluation chart

General Audio Understanding accuracy comparison across open-source and closed-source models

Speech Description

Fine-grained speech-style descriptions are evaluated across 13 descriptive dimensions using an LLM-as-judge protocol.

LLM-Judge Score ↑

Model evaluation chart

ASR

Summary of CER results across 12 ASR evaluation dimensions. Lower is better.

CER ↓

ModelOverallHealthDialectSingingNon-SpeechCode-SwitchCleanNoisyWhisperFar/NearMulti-SpkAgeSemantic
Paraformer-Large15.7722.1843.4532.344.9512.653.114.675.0217.4620.3314.967.14
GLM-ASR-Nano17.2924.4922.3951.954.6511.883.685.024.9427.5128.0217.197.32
Fun-ASR-Nano12.0421.997.8019.354.7611.232.983.463.7818.3819.8214.956.08
SenseVoice-Small14.5024.048.8923.794.9213.904.134.935.5726.6624.0617.637.55
Kimi-Audio-7B-Instruct14.1221.1129.3421.764.6816.382.202.152.6621.0220.6116.746.12
Qwen2.5-Omni-3B15.2624.6533.8724.245.5411.662.763.564.3222.1522.9115.177.24
Qwen2.5-Omni-7B15.0523.8531.9122.694.5612.972.523.163.6425.3821.0116.136.78
Qwen3-Omni-30B-A3B-Instruct11.3920.7315.6316.014.7311.302.232.471.9017.0818.1511.465.74
MOSS-Audio-4B-Instruct11.5821.1111.8410.794.0110.113.113.723.2918.4820.3315.098.15
MOSS-Audio-8B-Instruct11.3019.188.769.814.3110.182.703.202.7524.0424.3615.267.69

Timestamped ASR

Timestamp alignment quality is measured by AAS on Chinese and English benchmarks. Lower is better.

AAS ↓

ModelAISHELL-1 (zh)LibriSpeech (en)
Qwen3-Omni-30B-A3B-Instruct833.66646.95
Gemini-3.1-Pro708.24871.19
MOSS-Audio-4B-Instruct76.96358.13
MOSS-Audio-8B-Instruct35.77131.61

Model Demos

Released Models

ModelAudio EncoderLLM BackboneTotal SizeHugging FaceModelScope
MOSS-Audio-4B-InstructMOSS-Audio-EncoderQwen3-4B~4.6BModelModel
MOSS-Audio-4B-ThinkingMOSS-Audio-EncoderQwen3-4B~4.6BModelModel
MOSS-Audio-8B-InstructMOSS-Audio-EncoderQwen3-8B~8.6BModelModel
MOSS-Audio-8B-ThinkingMOSS-Audio-EncoderQwen3-8B~8.6BModelModel