• Music Understanding
  • Lyrics Recognition
  • Music Captioning
  • Chord Analysis
  • Long-Form Music Q&A

MOSS-Music

An open-source large language model for music understanding. Built on the MOSS-Audio backbone and continually pretrained on large-scale music data, it supports lyrics ASR, music captioning, chord/key/rhythm analysis, structural segmentation, and long-form music Q&A. It goes beyond merely “hearing music” to understanding how music is organized.

Author
OpenMOSS Team
Version
8B Instruct / 8B Thinking
Open-Source License
Apache 2.0

Model Overview

Music is never just about “hearing the lyrics” or “recognizing sound events.” At the surface of a song are vocals, instruments, melody, and production texture; beneath them lie chords, harmonic progression, key, tempo, beat, section structure, and emotional tension. General-purpose audio models have often been better at recognizing “what sounds occurred.” MOSS-Music instead asks how those sounds form a song: What key is it in? Where does the chorus begin? How do the lyrics, vocals, instrumentation, and chords relate to one another?

MOSS-Music is jointly open-sourced by MOSI.AI, the OpenMOSS team, and the Shanghai Innovation Institute. It follows the modular architecture of MOSS-Audio, comprising a dedicated audio encoder, a modality adapter, and a large language model. Raw audio is first encoded into a continuous temporal representation at 12.5 Hz, then projected into the language model's embedding space. The model is then continually pretrained and instruction fine-tuned on large-scale music data, with particular emphasis on singing, lyrics, full-length songs, harmony, structure, and long-form understanding. This release includes two 8B models: MOSS-Music-8B-Instruct follows direct instructions for music captioning, Q&A, lyrics ASR, and automatic tagging; MOSS-Music-8B-Thinking supports explicit reasoning for tasks requiring extended analysis of harmony, structure, style, and long-form music.

Capabilities

Lyrics ASR with Timestamps

Recognizes sung Chinese and English lyrics despite accompaniment and supports sentence-level and word-level timestamps. On three singing benchmarks, MUSDB18, MIR-1K, and Opencpop, the Thinking variant achieves an average WER / CER of 15.88%.

Music Captioning and Automatic Tagging

Generates natural-language descriptions of style, mood, tempo, instrumentation, vocals, melody, harmony, structure, production texture, and use cases. It can also produce concise tags on request. Across nine dimensions scored by GPT-5.4, MOSS-Music leads on both MusicCaps and SDD, with especially strong gains in describing structure and song sections.

Harmony, Key, and Rhythm Analysis

Identifies key, beat, downbeats, BPM, and chord progressions. It supports chord transcription and timestamped chord transcription for harmonic analysis, accompaniment reference, and music education.

Structural Segmentation

Segments full-length songs into sections such as intro, verse, chorus, bridge, and outro, and provides section-boundary timestamps to capture long-range musical organization.

Instrument and Vocal Part Recognition

Identifies lead instruments, vocal types, solo or choral singing, gender, and vocal range.

Long-Form Music Q&A and Reasoning

Answers open-ended questions about complete works, with retrospective analysis of sections, lyrics, structure, emotional changes, and musical elements. The Thinking variant presents its analysis before giving a final answer.

Time-Aware Model Design

During pretraining, temporal tokens are inserted into audio-frame representations at fixed intervals, allowing the model to learn “what happened when” in a unified text-generation framework. DeepStack-style cross-layer feature injection separately projects features from early, intermediate, and final encoder layers into the first few language-model layers, preserving low-level details such as rhythmic transients and timbre. This supports timestamped lyrics ASR, beat and downbeat localization, and section-boundary detection without an additional alignment module.

Open-Source Data Pipeline

MOSS-Music-Data-Pipeline is also open-sourced. It covers the full workflow from raw audio to chat-formatted training samples: duration statistics, full-song MIR analysis and structure annotation, structure-based segmentation, segment-level lyrics ASR and key/chord analysis, metadata cleaning, and caption/query generation. JSONL sharding supports large-scale corpus construction on HPC, Kubernetes, or Ray clusters.

Evaluation Results

Music Q&A and Understanding (Accuracy ↑)

Avg is the mean across eight public music Q&A benchmarks: MMAU-music, MMAU-mini-music, MMAU-Pro-music, MMAR-music, MuChoMusic, Music-AVQA, GTZAN, and Medley-Solos-DB. The three NSynth single-note recognition tasks are listed separately and excluded from the average.

ModelMMAU-musicMMAU-miniMMAU-ProMMARMuChoMusicMusic-AVQAGTZANMedley-Solos-DBAvg
MOSS-Music-8B-Instruct79.3380.7871.0259.7089.3976.7893.5992.4280.38
Gemini-3.1-Pro71.6977.1873.0671.6479.5361.5186.3980.3475.17
MOSS-Music-8B-Thinking74.0977.7867.9850.2582.9068.9084.7887.4274.26
MusicFlamingo76.8376.3565.6048.6674.5873.6084.4590.8673.87
Audio-Flamingo-Next72.3972.0761.6445.2775.6262.9477.6891.4769.89
MiMo-Audio-7B-Instruct66.3672.9766.5045.7775.4057.0565.6793.8167.94
Step-Audio-R166.4675.0862.3450.7572.6257.9873.6782.4567.67
Qwen3-Omni65.7668.7766.2748.5478.7756.0580.1569.6566.75
Kimi-Audio-7B-Instruct47.9552.2559.1045.2770.1868.9039.5471.9856.90

Music Captioning (GPT-5.4 Judge, 1–5)

Scores are averaged across nine dimensions: style, mood, tempo, instrumentation, vocals, melody/harmony, structure, production, and context of use.

ModelMusicCaps AvgSDD Avg
MOSS-Music-8B-Thinking4.534.45
MOSS-Music-8B-Instruct4.364.58
Gemini-3.1-Pro4.424.48
MusicFlamingo4.214.26
Audio-Flamingo-Next4.104.16
Qwen3-Omni3.964.02
MiMo-Audio-7B-Instruct3.984.05
Step-Audio-R13.903.98
Kimi-Audio-7B-Instruct3.733.80

Lyrics ASR (WER / CER ↓)

MUSDB18 contains English pop songs with accompaniment (WER), MIR-1K contains Chinese karaoke clips with accompaniment (CER), and Opencpop contains clean Mandarin studio vocals (CER).

ModelMUSDB18 WERMIR-1K CEROpencpop CERAvg
MOSS-Music-8B-Thinking29.19%15.84%2.60%15.88%
MOSS-Music-8B-Instruct32.99%23.96%4.62%20.52%
Gemini-3.1-Pro-Preview26.25%36.37%6.00%22.87%
MusicFlamingo23.41%38.98%18.73%27.04%
Qwen3-Omni-30B-A3B-Instruct62.67%20.48%2.26%28.47%
MiMo-Audio-7B-Instruct94.16%23.34%6.77%41.42%
Kimi-Audio-7B-Instruct97.53%25.83%4.90%42.75%
Step-Audio-R181.67%48.03%4.15%44.62%
Audio-Flamingo-Next94.93%55.63%12.47%54.34%

Chord Transcription

Chord transcription and timestamped chord transcription are supported. Detailed benchmark results will be published in a future update.

Model Demos