Professor Xipeng Qiu, Founder of MOSI, Invited to Attend the 4th China AIGC Industry Summit and Deliver a Keynote Address on “MOSS Multimodal Models and Inference Optimization”

MOSS Founder Professor Xipeng Qiu Delivers a Keynote Address on “MOSS Multimodal Models and Inference Optimization”

On May 20, 2026, the fourth China AIGC Industry Summit, hosted by QbitAI, was held in Beijing. With “@Everyone, Let’s Put AI to Work Now” as its theme, the summit brought together nearly 20 leading experts and business leaders from industry and academia to explore future trends in generative AI, large model technologies, and real-world industry applications. More than 1,000 people attended in person, while the online livestream attracted nearly 10 million viewers, drawing extensive attention and coverage from major technology media outlets.

MOSS Founder Professor Xipeng Qiu Delivers a Keynote Address on “MOSS Multimodal Models and Inference Optimization”

Professor Xipeng Qiu, Distinguished Professor at Fudan University, Assistant to the Dean of the Shanghai Innovation Institute, and Founder of MOSS, was invited to attend the summit and deliver a keynote address titled “MOSS Multimodal Models and Inference Optimization,” sharing insights into cutting-edge technologies with industry peers.

MOSS Founder Professor Xipeng Qiu Delivers a Keynote Address on “MOSS Multimodal Models and Inference Optimization”

Professor Qiu noted that we will inevitably enter an era of general contextual intelligence, in which interaction will play a particularly important role. Future multimodal models will need to fully understand context while handling longer contexts, greater Token consumption, and growing demand for real-time inference. Key highlights from the talk are summarized below. In his presentation on “MOSS Multimodal Models and Inference Optimization,” Professor Qiu put forward several key perspectives on the capability boundaries of large multimodal models, the demands of real-time interaction, and the key technical directions for the future era of intelligence:

MOSS Founder Professor Xipeng Qiu Delivers a Keynote Address on “MOSS Multimodal Models and Inference Optimization”

AI Will Enter an Era of General Contextual Intelligence, with Interaction Playing a Key Role

More powerful AI systems will need a deep understanding of context in its broadest sense, while interaction between humans and AI will become an important expression of the value of intelligent systems. Multimodal understanding must extend beyond text to include signals such as vision and audio, enabling AI to understand the full context surrounding the user.

Challenges for Multimodal Models in Real-Time Interaction

Multimodal AI designed for real-time interaction must handle longer contexts and more complex dimensions of information while meeting the demands of high-performance, low-latency inference. This creates new challenges for model architecture and inference design. Multimodal tasks consume far more Tokens than traditional text or coding tasks, placing stricter requirements on both models and inference frameworks.

MOSS Founder Professor Xipeng Qiu Delivers a Keynote Address on “MOSS Multimodal Models and Inference Optimization”

Video Understanding and Temporal Reasoning Will Become Core Capabilities

Video understanding has unique advantages in information density and temporal logic, making it one of the core capabilities of future multimodal interaction systems.

Key Innovations Across the MOSS Model Family

MOSS-VL uses a cross-attention architecture to support continuous video input, allowing the language model to access dynamic visual information on demand for more natural interaction; MOSS-VL Goes Open Source: A Cross-Attention Architecture Drives a New Paradigm for Video Understanding, Adding Another Core Piece to the Open MOSS Multimodal Ecosystem

MOSS-Audio is designed to understand audio in a broader context, going beyond recognizing spoken content to interpret complex information such as scenes and emotions. It has reached the level of leading specialized models on tasks including ASR, speech captioning, and timestamped ASR; MOSS-Audio: One Model to Understand Every Sound

MOSS-TTS provides end-to-end capabilities covering speech synthesis, lightweight deployment, voice design, and real-time performance. Its audio tokenizer, built on a pure Transformer architecture, has been widely adopted and surpassed one million downloads after its open-source release. The MOSS-TTS Family Is Officially Released: True Production-Ready Coverage Across Use Cases—Voice Cloning, Long-Form Speech, Dialogue, Instruction Following, and Sound Effects

In his keynote, Professor Qiu emphasized that future intelligent interaction systems should seamlessly integrate visual understanding, speech understanding, and speech output into end-to-end models designed for contextual interaction. This is a critical step toward bringing AI into broader applications.