MOSS Makes Its Official Debut at the 2026 World Artificial Intelligence Conference (WAIC)

As an innovative large-model company focused on next-generation contextual interactive intelligence, MOSS showcased a range of cutting-edge technological achievements, including real-time video understanding, complex audio transcription, speech and audio/video generation, and real-time interaction.

From July 17 to 20, 2026, the 2026 World Artificial Intelligence Conference and the High-Level Conference on Global AI Governance were held in Shanghai. MOSS, an innovative large-model company focused on next-generation contextual interactive intelligence, made its debut at WAIC. At booth H1-C1234 in the Shanghai World Expo Exhibition and Convention Center, MOSS showcased its latest advances in real-time video understanding, complex audio transcription, speech and audio/video generation, and real-time interaction.

MOSS WAIC Booth

Professor Xipeng Qiu of Fudan University, Full-Time Mentor at the Shanghai Innovation Institute, and Chief Scientist of MOSS, was invited to attend the roundtable at the main forum of the 2026 World Artificial Intelligence Conference. He joined Omar M. Yaghi, winner of the 2025 Nobel Prize in Chemistry, Gilles Brassard, winner of the 2025 Turing Award and a pioneer in quantum technology, and other leading international scientists on stage to exchange views on AI-driven changes in scientific research paradigms, the evolution of next-generation models, and the development of artificial intelligence and quantum technology.

New Paradigms for Scientific Research Forum

At the forum, Professor Qiu stated that AI is evolving from an auxiliary tool for scientific research into an important participant in scientific innovation. In the future, a closed-loop research and development system in which “AI autonomously iterates AI” will emerge, continuously shortening model development cycles and accelerating the development of next-generation foundation models.

Before the conference opened, MOSS and the OpenMOSS team released several models and products in succession: on July 14, they open-sourced the real-time visual understanding model MOSS-VL-Realtime; on July 10, they opened the beta testing of the one-stop AIGC creation platform Mossland and the MOSS API open platform; and on July 9, they open-sourced the multi-speaker long-form audio transcription model MOSS-Transcribe-Diarize-0.9B.

Overview of the MOSS WAIC Booth

In addition to demonstrating the strong performance and user experience of its core models, MOSS turned these capabilities into an interactive experience that anyone could participate in at this year’s WAIC: say a sentence to light up a Sound Babel Tower; make a phone call to generate a dubbed segment or an AI podcast; or speak into a microphone to receive a “voice personality card” that records speech rate, rhythm, tone, and emotion.

Together, these experiences reflect MOSS’s view of next-generation AI interaction: when models enter the real world, the key is not only to recognize sounds and images, but also to understand the context in which this information occurs and use that understanding to determine when and how to respond. This is MOSS’s answer to “contextual intelligence.”

Starting with the Sound Babel Tower: Making Contextual Intelligence Visible

Visitors can speak languages from different countries or dialects from different regions. After the system identifies the language, accent, and expressive characteristics, it lights up a node on the tower. As more people participate, the Babel Tower is completed layer by layer and ultimately illuminated by voices from different linguistic backgrounds.

MOSS Contextual Experience Zone

The ancient parable of the Tower of Babel tells a story of linguistic divergence and the breakdown of understanding. MOSS turns this story on its head: language once prevented people from understanding one another, while AI today has the opportunity to overcome differences in language, accents, and communication habits, taking communication beyond “correct translation” toward “accurate understanding.”

This installation is powered by the ability of MOSS-Transcribe and MOSS-Audio to recognize and understand complex sounds. The models must not only hear what is being said clearly, but also understand tone, emotion, rhythm, and context as fully as possible. In scenarios such as cross-language communication and global customer service, this non-textual information often directly affects the interpretation of intent.

The booth also featured the Mossland Creation Phone Booth and the VBTI (Voice-Based Type Indicator) voice personality test. The phone booth is powered by MOSS’s self-developed MOSS-TTS Family speech synthesis model: visitors can speak an idea as if making a phone call and use it to dub a film or television clip, generate an AI podcast, or receive a “voice preservation ticket” recording their voice at that moment. When visitors speak into the microphone, the system generates a “voice personality card” showing their speech rate, rhythm, tone, and emotions.

The Sound Babel Tower senses different types of sound information, VBTI further interprets the emotions and expressive states within that information, and the creation phone booth transforms this understanding into content such as dubbing and podcasts. Together, the three experiences present a contextual intelligence pathway from perception and understanding to generation and interaction. From real-time video to multi-speaker long-form audio, the system continuously understands unfolding contexts.

Everything Enters the Context, Intelligence Is Reborn

Behind these on-site experiences is the MOSS family of multimodal foundation models that MOSS is building. At this year’s WAIC, MOSS brought its MOSS-TTS speech synthesis model, MOSS-Transcribe-Diarize multi-speaker transcription model, and the newly released MOSS-VL-Realtime real-time visual understanding model together to the exhibition.

Inside the booth, a model wall approximately six meters long was open for developers and visitors to try: MOSS-VL-Realtime continuously understood and explained the scenes around it, MOSS-TTS generated audio works based on different requirements, and MOSS-Transcribe-Diarize distinguished different speakers and transcribed speech in the noisy exhibition environment. Rather than simply displaying model parameters, these capabilities ran directly in a real, complex, and constantly changing environment.

MOSS-VL-Realtime

It is designed not for a recording that has already ended, but for a video stream that is continuously unfolding. Traditional video understanding models typically need to watch an entire video before answering questions based on its content. However, the scenes in livestreams, sporting events, real-time monitoring, and smart hardware do not pause, and information that appears later may change an earlier judgment at any time.

MOSS-VL-Realtime can answer questions while video continues to stream in. When key information has not yet appeared, the model continues observing and remains silent. When the scene changes, it can also promptly correct an existing answer. It supports video understanding at the scale of hours and can increase the sampling frequency for rapidly changing scenes, reducing the chance of missing key actions.

In comprehensive evaluations on streaming video understanding benchmarks including OVOBench, OmniMMI, and ProactiveVideoQA, MOSS-VL-Realtime achieved the best performance among open-source models.

From Models to Real-World Use: Completing the Contextual Intelligence Loop

MOSS-Transcribe-Diarize

On the audio side, MOSS-Transcribe-Diarize-0.9B is designed for complex long-form audio scenarios such as multi-speaker meetings, interviews, and podcasts. It needs to identify not only “what was said,” but also “who said it” and “when it was said.”

Previous approaches typically split automatic speech recognition, speaker diarization, and temporal alignment into multiple stages. Once an earlier stage made an error, that error could continue to propagate. MOSS-Transcribe-Diarize processes text, speaker, and timing information within a single model and can directly output structured text with speaker labels and timestamps.

With only 0.9B parameters, the model is suitable for edge deployment. It can process approximately 90 minutes of multi-speaker audio in a single run, distinguish different speakers, and support hotword enhancement for names, brand names, and technical terms. In the AISHELL-4 public benchmark for Chinese multi-speaker meetings, the model achieved the best results across all three core metrics: transcription accuracy, speaker diarization, and stability.

Video and audio may appear to be two different technical tracks, but they address the same underlying problem: the model needs to track the relationships among events, people, and time as information continuously arrives. A video model must understand how a scene changes, while an audio model must continuously associate speakers, content, and timing.


MOSS was incubated by the Shanghai Innovation Institute and the OpenMOSS team at Fudan University. Professor Xipeng Qiu of Fudan University and Full-Time Mentor at the Shanghai Innovation Institute serves as its Chief Scientist. Starting with the MOSS language model in 2023, the team has gradually expanded its capabilities to the understanding and generation of multiple modalities, including speech, audio, and vision, as well as video and multimodal generation.

MOSS Exhibition Booth

MOSS is building the MOSS family of models around perception and understanding, content generation, and real-time interaction. This approach does not simply combine speech recognition, video understanding, and content generation. Instead, it aims to preserve the tone, emotion, rhythm, identity, and on-site state that are easily lost during transcription in real-world communication, enabling the model to understand sound, images, and time within a complete context.

These capabilities are now beginning to move from models and benchmarks into real-world use. The MOSS-TTS Family has surpassed 2 million cumulative downloads across open-source communities such as Hugging Face. Following its open-source release, MOSS-Transcribe-Diarize-0.9B reached first place on the Hugging Face Audio-Text-to-Text Trending list. OpenMOSS-related open-source projects have also received more than 10,000 Stars on GitHub.

MOSS Gaining Momentum

During its closed beta, Mossland, a platform for creators, attracted more than 100,000 active creators. It brings capabilities such as voice cloning, timbre design, video dubbing, and AI podcast creation into a single creative workflow. MOSS API further makes capabilities including speech generation, long-form audio transcription, and audio/video understanding available to developers and enterprises.

Foundation models enter real-world scenarios through products, while users’ actual usage continuously provides new data and feedback to the models, helping the team optimize understanding, generation, and interaction capabilities. In this way, MOSS has gradually formed a complete loop from foundation models to product applications, user feedback, and back to model iteration.

When large models enter agents, robots, smart cabins, companion devices, and other real-world terminals, they will no longer face a neatly organized piece of text. Instead, they will encounter continuously unfolding sounds, images, actions, relationships among people, and changes in the environment. AI needs to understand the present and adjust its judgments and responses promptly as the context changes.


From understanding video in real time, to clearly hearing multi-speaker long-form audio, to understanding tone, emotion, and on-site states. MOSS hopes to connect AI, humans, and the physical world through contextual intelligence, enabling models to move from processing isolated information to understanding continuously unfolding reality. People will no longer need to organize complex intentions into a single precise instruction. AI can understand people’s states, goals, and changes from the context as it unfolds, participate more naturally in real-world collaboration, and better understand and serve people.

Front View of the MOSS Exhibition Booth