Today’s AI can already read text, recognize images, understand sounds, and even generate coherent videos. But when a model has language, vision, and hearing capabilities at the same time, does it truly understand the world?

Professor Xipeng Qiu of Fudan University, Head of the OpenMOSS Project, and Chief Scientist of MOSS, joined Omar M. Yaghi, winner of the 2025 Nobel Prize in Chemistry, Gilles Brassard, winner of the 2025 Turing Award and a pioneer in quantum information technology, and CAS Academician E Weinan for a discussion at the main forum of the World Artificial Intelligence Conference on July 17.
Professor Qiu shared his views on multimodality, real-world understanding, and next-generation AI paradigms. In his view, the future of AI will not be defined solely by adding more modalities such as images, video, and sound, nor by continuously expanding the model’s context window. The more fundamental question is:
Can AI understand what people mean in a specific situation and the real context formed by the complex environment around them?
This points to a capability that moves from “processing information” to “understanding context.”
Language Remains a Key Foundation for Intelligent Understanding of the World
There are two different perspectives in academia and industry regarding next-generation AI.
One view holds that language models will remain the core of future AI. Different modalities, including images, sounds, and video, can all be converted into a unified representation that models can process, with language models ultimately responsible for understanding, reasoning, and generation.
Another view is that when AI moves from digital spaces into the physical world and enters scenarios such as robotics and world models, multimodality will no longer be a simple extension of language models, but will become an intrinsic part of intelligent systems.

Faced with these two directions, Professor Qiu believes that intelligence will still need to rely on language, which is an important channel through which humans understand and make sense of the world.
Human cognition, description, and organization of knowledge about the world depend to a great extent on language. Language is not only a tool for communication, but also an important way for humans to construct abstract concepts and express complex relationships. This does not mean, however, that the problem is solved once images, sounds, and videos are converted into language. The true challenge of multimodality is how to effectively align information from different sources with language, allowing models not only to recognize individual pieces of information, but also to understand the relationships among them. Identifying what appears in an image, what is said in an audio clip, or what happens in a video is only the first step toward understanding the world. Models must go further and determine:
Why did these things happen?
What connections exist among the different pieces of information?
What kind of real-world state do they collectively form?
The value of multimodality lies not in giving models more input channels, but in helping them develop a more complete understanding of the world.
The Context AI Needs to Understand Is More Than a Chat Window
Over the past few years, large models have continued to expand their context lengths. From a few thousand tokens to hundreds of thousands or even millions of tokens, models can read and process more and more information in a single pass. But the context that Professor Qiu emphasizes is not simply about how long the context is or how large the window is. It points to the much more complex “situations” of the real world.

For example, when a robot enters a real space and interacts with people, it does not face a neatly organized piece of text. It must process all of the following at the same time:
What the speaker is expressing;
What tone the speaker is using;
The setting in which the conversation takes place;
The people and objects nearby;
The relationships among the people involved;
What happened previously;
What may happen next.
The same phrase, “It’s okay,” might be used to comfort someone or to conceal disappointment. The same sentence, “You can go back first,” may have entirely different meanings in a workplace, a family setting, or an intimate relationship. Humans can often understand these differences quickly based on tone, facial expressions, the environment, and relationships. For current models, however, this remains a difficult task. Professor Qiu pointed out that models’ understanding of the world is still far from sufficient. A model may already be able to identify a person, an object, or a sentence, but that does not necessarily mean it understands what the person truly intends to express in a specific situation. Nor can it fully grasp how the surrounding environment affects the interaction.
Therefore, next-generation AI needs more than longer text memory. It needs a more complete ability to model real-world situations. Discussions of context should not stop at the length of the text window, but should extend to a deeper understanding of complex real-world contexts.
Next-Generation AI May Be Designed with AI’s Participation
When discussing the development of artificial intelligence over the next five years, Professor Qiu first emphasized that predicting the future of technology is extremely difficult.
Looking back over the past few years, many technological changes that seem natural today were not accurately anticipated at the time. Some directions may have been overestimated, while other changes that ultimately had a profound impact may have been underestimated for a long time. Technological breakthroughs do not necessarily follow paths defined in advance. Sometimes, goals that people deliberately pursue may not be achieved as expected, while technologies that seem to attract little attention today may gain new opportunities after continued accumulation.

Among the many possible futures, Professor Qiu pointed to one direction worth watching:
AI for AI: Our Next-generation AI May Itself Be Designed by AI.
As coding model and agent model capabilities continue to advance, Professor Qiu believes that these capabilities could first be developed further, after which AI could be invited to participate in the design of future multimodal models. This would allow researchers to explore whether better model structures and technical paradigms exist. This does not mean that next-generation models can already be developed autonomously by AI. Instead, it offers a new possibility: AI may not only be applied to other tasks, but may gradually participate in the research and design of artificial intelligence itself.
Genuine Paradigm Shifts Usually Do Not Happen Overnight
Professor Qiu said, “In 2023, GPT looked like a major paradigm shift to everyone, but it also emerged step by step.”
Before ChatGPT, ideas such as sequence-to-sequence, sequence generation, and unified natural language processing tasks had already undergone years of exploration. Researchers first recognized the potential of a technical approach, then continuously optimized model structures, training methods, data, and engineering systems. It was only when model capabilities crossed a certain threshold that a new form of interaction and product experience truly emerged.
Therefore, future multimodal intelligence may not need to completely overturn today’s large-model systems. It may develop gradually from existing language models, coding models, and agent models, driven by real-world usage requirements, without necessarily requiring a complete reinvention of the existing system. A new paradigm is often not designed all at once. It is more likely to grow gradually through continuous use, ongoing feedback, and long-term iteration.
From Understanding Language to Understanding the World as It Unfolds
Returning to MOSS’s technical practice, these discussions also provide a perspective for understanding “contextual intelligence.”
Traditional AI products typically receive a clearly expressed instruction. A user enters a question, and the model generates an answer.
But when AI truly enters meetings, livestreams, education, content creation, and real-time interaction scenarios, it no longer faces a static question, but an ongoing process. Multiple people may speak at the same time, subjects in the frame may keep moving, and the user’s emotions may change...
The most important information may appear only briefly within a long piece of content, and the model must also determine when it should respond and when it should continue observing.

Take sound as an example. The information conveyed by humans has never existed only in the words themselves. The speaker’s identity, speech rate, rhythm, pauses, timbre, and emotions are equally important parts of communication. The same sentence may convey entirely different intentions when spoken by different people or in different relationships and environments. Truly natural audio intelligence therefore requires more than enabling AI to generate realistic voices. It also needs to understand:
Who is speaking;
What emotional state the speaker is in;
What kind of relationship and situation the statement occurs in;
What tone should be used in the response.

In this direction, MOSS continues to advance the development of the MOSS family of multimodal models, covering capabilities such as visual understanding, audio understanding, speech generation, and real-time interaction. MOSS-VL-Realtime is designed to understand continuously incoming video streams in real time. When key information has not yet appeared, the model can continue observing instead of rushing to a conclusion. MOSS-Transcribe-Diarize is designed for multi-speaker long-form audio scenarios such as meetings, interviews, and podcasts, processing speech content, speakers, and timing information within a single model.
These capabilities are not simply created by combining multiple models. Together, they point toward one goal:
Enable AI to understand the constantly changing relationships among sounds, images, people, and environments along a continuous timeline.