In the second Multilingual Conversational Speech Language Model Challenge (MLC-SLM), our model MOSS-Transcribe-Diarize-Pro took first place in Task 1 with a tcpMER score of 14.93!

The real difficulty of this evaluation was that the model was not processing standard, clear single-speaker readings. Instead, it had to handle 14 languages, accents from multiple regions, and real multi-speaker conversations.
These included: English (en), French (fr), German (de), Italian (it), Portuguese (pt), Spanish (es), Japanese (jp), Korean (ko), Russian (ru), Thai (th), Vietnamese (vi), Tagalog (tl), Urdu (ur), and Turkish (tr).
The accent diversity was also considerable: British English, American English, Australian English, Indian English, Philippine English, Canadian French, Mexican Spanish, and Brazilian Portuguese.
MLC-SLM is a satellite event of Interspeech 2026, attracting 110 teams from industry and academia worldwide.
The competition provided nearly 1,500 hours of conversational speech data, including approximately 500 hours of English. With different languages, regional accents, multiple speakers taking turns, and frequent interruptions, the model had to do more than recognize “what was said”—it also had to continuously determine “who was speaking.”
Under this unified evaluation, MOSS-Transcribe-Diarize-Pro took first place!
It completes speech recognition, speaker attribution, and timestamp prediction in a single inference pass, directly outputting structured transcription results without the need to combine multiple additional modules.
For different use cases, we also offer two open-source transcription models: MOSS-Transcribe-Diarize-0.9B for lightweight deployment, and MOSS-Transcribe for complex English scenarios, with dedicated optimizations for real-world conditions including standard English, diverse accents, and soft-spoken whispers.
From competition benchmarks to real-world business applications, we will continue making multi-speaker transcription more accurate, reliable, and user-friendly.