MOSS Appears at ChinaJoy, CEO Li Shimin Speaks on “Connecting AI, People, and the Physical World”

MOSS Co-founder and CEO Shimin Li Speaks on “Connecting AI, Humans, and the Physical World”

At the ChinaJoy AI Future Ecosystem Conference held on July 31, Shimin Li, Co-founder and CEO of MOSS, shared his views on next-generation human–computer interaction around the theme of “Connecting AI, Humans, and the Physical World”:

As interaction moves from text, PCs, and smartphones to voice, vision, spatial signals, and environmental signals, models need to understand more than a single instruction. They need to understand real-world contexts that are longer, more complex, and more diverse.

With contextual intelligence at its core, MOSS is building the MOSS family of large models for the real world.

MOSS CEO Li Shimin Delivers a Keynote on “Contextual Intelligence”

In speech, MOSS-Audio focuses on contextual understanding of sound and audio. MOSS-Transcribe-Diarize focuses on noisy, overlapping, and multi-speaker scenarios, enabling models to identify more accurately who said what and when. MOSS-TTS Family covers natural speech generation, voice cloning, long-form text, multi-speaker dialogue, real-time streaming speech, and voice design.

In vision, MOSS-VL-Realtime can watch, understand, and respond within continuous video streams, enabling more timely proactive interaction.

From sound as an entry point to full-modal understanding and generation;

From open models to MOSS Create and MossAPI;

MOSS is enabling AI to do more than recognize information. It can understand context and naturally connect people with the world.