
On May 20, 2026, the fourth China AIGC Industry Summit, hosted by QbitAI, was held in Beijing. With “@Everyone, Let’s Put AI to Work Now” as its theme, the summit brought together nearly 20 leading experts and business leaders from industry and academia to explore future trends in generative AI, large model technologies, and real-world industry applications. More than 1,000 people attended in person, while the online livestream attracted nearly 10 million viewers, drawing extensive attention and coverage from major technology media outlets.

Professor Xipeng Qiu, Distinguished Professor at Fudan University, Assistant to the Dean of the Shanghai Innovation Institute, and Founder of MOSS, was invited to attend the summit and deliver a keynote address titled “MOSS Multimodal Models and Inference Optimization,” sharing insights into cutting-edge technologies with industry peers.

Professor Qiu noted that we will inevitably enter an era of general contextual intelligence, in which interaction will play a particularly important role. Future multimodal models will need to fully understand context while handling longer contexts, greater Token consumption, and growing demand for real-time inference. Key highlights from the talk are summarized below. In his presentation on “MOSS Multimodal Models and Inference Optimization,” Professor Qiu put forward several key perspectives on the capability boundaries of large multimodal models, the demands of real-time interaction, and the key technical directions for the future era of intelligence:

AI Will Enter an Era of General Contextual Intelligence, with Interaction Playing a Key Role
More powerful AI systems will need a deep understanding of context in its broadest sense, while interaction between humans and AI will become an important expression of the value of intelligent systems. Multimodal understanding must extend beyond text to include signals such as vision and audio, enabling AI to understand the full context surrounding the user.
Challenges for Multimodal Models in Real-Time Interaction
Multimodal AI designed for real-time interaction must handle longer contexts and more complex dimensions of information while meeting the demands of high-performance, low-latency inference. This creates new challenges for model architecture and inference design. Multimodal tasks consume far more Tokens than traditional text or coding tasks, placing stricter requirements on both models and inference frameworks.

Video Understanding and Temporal Reasoning Will Become Core Capabilities
Video understanding has unique advantages in information density and temporal logic, making it one of the core capabilities of future multimodal interaction systems.
Key Innovations Across the MOSS Model Family
MOSS-VL uses a cross-attention architecture to support continuous video input, allowing the language model to access dynamic visual information on demand for more natural interaction; MOSS-VL Goes Open Source: A Cross-Attention Architecture Drives a New Paradigm for Video Understanding, Adding Another Core Piece to the Open MOSS Multimodal Ecosystem
MOSS-Audio is designed to understand audio in a broader context, going beyond recognizing spoken content to interpret complex information such as scenes and emotions. It has reached the level of leading specialized models on tasks including ASR, speech captioning, and timestamped ASR; MOSS-Audio: One Model to Understand Every Sound
MOSS-TTS provides end-to-end capabilities covering speech synthesis, lightweight deployment, voice design, and real-time performance. Its audio tokenizer, built on a pure Transformer architecture, has been widely adopted and surpassed one million downloads after its open-source release. The MOSS-TTS Family Is Officially Released: True Production-Ready Coverage Across Use Cases—Voice Cloning, Long-Form Speech, Dialogue, Instruction Following, and Sound Effects
In his keynote, Professor Qiu emphasized that future intelligent interaction systems should seamlessly integrate visual understanding, speech understanding, and speech output into end-to-end models designed for contextual interaction. This is a critical step toward bringing AI into broader applications.