Model Overview
Music is never just about “hearing the lyrics” or “recognizing sound events.” At the surface of a song are vocals, instruments, melody, and production texture; beneath them lie chords, harmonic progression, key, tempo, beat, section structure, and emotional tension. General-purpose audio models have often been better at recognizing “what sounds occurred.” MOSS-Music instead asks how those sounds form a song: What key is it in? Where does the chorus begin? How do the lyrics, vocals, instrumentation, and chords relate to one another?
MOSS-Music is jointly open-sourced by MOSI.AI, the OpenMOSS team, and the Shanghai Innovation Institute. It follows the modular architecture of MOSS-Audio, comprising a dedicated audio encoder, a modality adapter, and a large language model. Raw audio is first encoded into a continuous temporal representation at 12.5 Hz, then projected into the language model's embedding space. The model is then continually pretrained and instruction fine-tuned on large-scale music data, with particular emphasis on singing, lyrics, full-length songs, harmony, structure, and long-form understanding. This release includes two 8B models: MOSS-Music-8B-Instruct follows direct instructions for music captioning, Q&A, lyrics ASR, and automatic tagging; MOSS-Music-8B-Thinking supports explicit reasoning for tasks requiring extended analysis of harmony, structure, style, and long-form music.
Capabilities
Lyrics ASR with Timestamps
Recognizes sung Chinese and English lyrics despite accompaniment and supports sentence-level and word-level timestamps. On three singing benchmarks, MUSDB18, MIR-1K, and Opencpop, the Thinking variant achieves an average WER / CER of 15.88%.
Music Captioning and Automatic Tagging
Generates natural-language descriptions of style, mood, tempo, instrumentation, vocals, melody, harmony, structure, production texture, and use cases. It can also produce concise tags on request. Across nine dimensions scored by GPT-5.4, MOSS-Music leads on both MusicCaps and SDD, with especially strong gains in describing structure and song sections.
Harmony, Key, and Rhythm Analysis
Identifies key, beat, downbeats, BPM, and chord progressions. It supports chord transcription and timestamped chord transcription for harmonic analysis, accompaniment reference, and music education.
Structural Segmentation
Segments full-length songs into sections such as intro, verse, chorus, bridge, and outro, and provides section-boundary timestamps to capture long-range musical organization.
Instrument and Vocal Part Recognition
Identifies lead instruments, vocal types, solo or choral singing, gender, and vocal range.
Long-Form Music Q&A and Reasoning
Answers open-ended questions about complete works, with retrospective analysis of sections, lyrics, structure, emotional changes, and musical elements. The Thinking variant presents its analysis before giving a final answer.
Time-Aware Model Design
During pretraining, temporal tokens are inserted into audio-frame representations at fixed intervals, allowing the model to learn “what happened when” in a unified text-generation framework. DeepStack-style cross-layer feature injection separately projects features from early, intermediate, and final encoder layers into the first few language-model layers, preserving low-level details such as rhythmic transients and timbre. This supports timestamped lyrics ASR, beat and downbeat localization, and section-boundary detection without an additional alignment module.
Open-Source Data Pipeline
MOSS-Music-Data-Pipeline is also open-sourced. It covers the full workflow from raw audio to chat-formatted training samples: duration statistics, full-song MIR analysis and structure annotation, structure-based segmentation, segment-level lyrics ASR and key/chord analysis, metadata cleaning, and caption/query generation. JSONL sharding supports large-scale corpus construction on HPC, Kubernetes, or Ray clusters.
Evaluation Results
Music Q&A and Understanding (Accuracy ↑)
Avg is the mean across eight public music Q&A benchmarks: MMAU-music, MMAU-mini-music, MMAU-Pro-music, MMAR-music, MuChoMusic, Music-AVQA, GTZAN, and Medley-Solos-DB. The three NSynth single-note recognition tasks are listed separately and excluded from the average.
| Model | MMAU-music | MMAU-mini | MMAU-Pro | MMAR | MuChoMusic | Music-AVQA | GTZAN | Medley-Solos-DB | Avg |
|---|---|---|---|---|---|---|---|---|---|
| MOSS-Music-8B-Instruct | 79.33 | 80.78 | 71.02 | 59.70 | 89.39 | 76.78 | 93.59 | 92.42 | 80.38 |
| Gemini-3.1-Pro | 71.69 | 77.18 | 73.06 | 71.64 | 79.53 | 61.51 | 86.39 | 80.34 | 75.17 |
| MOSS-Music-8B-Thinking | 74.09 | 77.78 | 67.98 | 50.25 | 82.90 | 68.90 | 84.78 | 87.42 | 74.26 |
| MusicFlamingo | 76.83 | 76.35 | 65.60 | 48.66 | 74.58 | 73.60 | 84.45 | 90.86 | 73.87 |
| Audio-Flamingo-Next | 72.39 | 72.07 | 61.64 | 45.27 | 75.62 | 62.94 | 77.68 | 91.47 | 69.89 |
| MiMo-Audio-7B-Instruct | 66.36 | 72.97 | 66.50 | 45.77 | 75.40 | 57.05 | 65.67 | 93.81 | 67.94 |
| Step-Audio-R1 | 66.46 | 75.08 | 62.34 | 50.75 | 72.62 | 57.98 | 73.67 | 82.45 | 67.67 |
| Qwen3-Omni | 65.76 | 68.77 | 66.27 | 48.54 | 78.77 | 56.05 | 80.15 | 69.65 | 66.75 |
| Kimi-Audio-7B-Instruct | 47.95 | 52.25 | 59.10 | 45.27 | 70.18 | 68.90 | 39.54 | 71.98 | 56.90 |
Music Captioning (GPT-5.4 Judge, 1–5)
Scores are averaged across nine dimensions: style, mood, tempo, instrumentation, vocals, melody/harmony, structure, production, and context of use.
| Model | MusicCaps Avg | SDD Avg |
|---|---|---|
| MOSS-Music-8B-Thinking | 4.53 | 4.45 |
| MOSS-Music-8B-Instruct | 4.36 | 4.58 |
| Gemini-3.1-Pro | 4.42 | 4.48 |
| MusicFlamingo | 4.21 | 4.26 |
| Audio-Flamingo-Next | 4.10 | 4.16 |
| Qwen3-Omni | 3.96 | 4.02 |
| MiMo-Audio-7B-Instruct | 3.98 | 4.05 |
| Step-Audio-R1 | 3.90 | 3.98 |
| Kimi-Audio-7B-Instruct | 3.73 | 3.80 |
Lyrics ASR (WER / CER ↓)
MUSDB18 contains English pop songs with accompaniment (WER), MIR-1K contains Chinese karaoke clips with accompaniment (CER), and Opencpop contains clean Mandarin studio vocals (CER).
| Model | MUSDB18 WER | MIR-1K CER | Opencpop CER | Avg |
|---|---|---|---|---|
| MOSS-Music-8B-Thinking | 29.19% | 15.84% | 2.60% | 15.88% |
| MOSS-Music-8B-Instruct | 32.99% | 23.96% | 4.62% | 20.52% |
| Gemini-3.1-Pro-Preview | 26.25% | 36.37% | 6.00% | 22.87% |
| MusicFlamingo | 23.41% | 38.98% | 18.73% | 27.04% |
| Qwen3-Omni-30B-A3B-Instruct | 62.67% | 20.48% | 2.26% | 28.47% |
| MiMo-Audio-7B-Instruct | 94.16% | 23.34% | 6.77% | 41.42% |
| Kimi-Audio-7B-Instruct | 97.53% | 25.83% | 4.90% | 42.75% |
| Step-Audio-R1 | 81.67% | 48.03% | 4.15% | 44.62% |
| Audio-Flamingo-Next | 94.93% | 55.63% | 12.47% | 54.34% |
Chord Transcription
Chord transcription and timestamped chord transcription are supported. Detailed benchmark results will be published in a future update.