Professor Xipeng Qiu explained that MOSS is designed for next-generation contextual intelligent interaction. It integrates visual, auditory, and generative capabilities across video understanding, audio understanding, speech dialogue, and audio/video generation. Through continuous perception, anytime interaction, dynamic error correction, and proactive responses, MOSS aims to help AI better understand real-world contexts and communicate naturally with people.

The evolution of model capabilities depends on the coordination of algorithms and computing power. MOSS maintains long-term collaboration with the Huawei Xiaoqiaoling team and the Ascend Computing Ecosystem team, advancing model research and applications through on-site support, training and inference adaptation, and joint performance optimization. In particular, MOSS’s first-packet latency for speech interaction has been reduced to 29 ms. Multi-speaker transcription, edge-cloud speech dialogue, and synchronized audio/video generation have all achieved leading performance among open-source solutions.
Open source has also brought these achievements into the hands of a broader developer community. As of this presentation, the MOSS multimodal model series has surpassed 7 million cumulative downloads on Hugging Face and gained support from open-source communities such as SGLang and vLLM.

Professor Qiu pointed out that AI is evolving from isolated tools into intelligent infrastructure for scientific research. Around model innovation and research applications, MOSS and Ascend are further co-building an open-source innovation ecosystem by advancing talent development through hands-on MOSS courses and developer training, improving community tutorials, evaluations, and application examples, and conducting joint research and development on model training and inference as well as AI for Science.
Looking ahead, MOSS will continue advancing research and development in contextual intelligence and multimodal foundation models, deepening collaboration with Ascend and open-source communities, and enabling model perception and interaction capabilities to be validated and applied across more real-world scenarios.
