MOSS’s Cheng Qinyuan: As AI Moves from the Digital World into the Physical World, Interaction Technology Will Be Redefined

Recently, Qinyuan Cheng, CTO of MOSS, accepted an exclusive interview with The Paper Technology. He shared his perspectives and practical insights on next-generation human–computer interaction, focusing on the technical barriers, data challenges, and commercialization hurdles facing speech models.

MOSS CTO Cheng Qinyuan in an exclusive interview with The Paper Technology

As AI moves beyond screens and code and gradually enters the real physical world, traditional human–computer interaction paradigms are bound to undergo a new transformation. Multimodality is the core gateway through which AI connects with the physical world, while speech is a crucial medium for human–computer interaction.


The following is an excerpt from the conversation. Click the audio link to listen to the full conversation as a podcast.

From PhD Student to CTO of a Speech Model Company

The Paper Technology: You started your own company before completing your PhD?

Qinyuan Cheng: Yes. I spent the last two years of my PhD building a company. The emergence of large models has changed the way PhD students and graduate students approach their work. Before large models, people might have focused on narrowly defined academic papers. After large models emerged, more outstanding PhD students began wanting to build projects with greater impact. Such projects are often reflected in strong visibility within the open-source community or performance approaching that of commercial models. These things are closely related to entrepreneurship. Generally speaking, startup-level resources and organization are required to accomplish something truly impactful.

The Paper Technology: MOSS is a language model, but after starting the company, you initially focused on speech models. Why did this shift happen?

Qinyuan Cheng: We are still working on language models as well. In the laboratory, our current focus for language models is mainly on post-training.

We decided to work on multimodality because we believe it is the inevitable next step after language models. Once you have a relatively capable agent that can perform well online or in a digital environment, it will eventually enter the real world. For example, it may operate through an embodied robot, or remain on a computer while interacting more extensively with the real world. In that case, multimodality is unavoidable—or, in other words, it becomes a necessary interface. We believe this transition is bound to happen.

Another reason is that our technical approach allows us to reuse much of the experience we gained from working on large models. For example, our speech models are still primarily based on discretization.

The Paper Technology: What aspects of speech models cannot simply be replicated from language models?

Qinyuan Cheng: As everyone knows, large models currently use tokens to represent their content. Speech differs from text in that converting text into tokens is relatively straightforward, while speech contains information that has not been abstracted particularly well, such as human emotion and expressiveness. This raises the question of how to encode information that does not appear in text or is not readily represented by text.

I would call these scientific problems because they do not have particularly standardized engineering solutions. They may require conducting experiments, proposing hypotheses, and validating them before arriving at a reasonably good result and then scaling it up.

How Speech Models Have Evolved from the Past to the Future

The Paper Technology: We usually do not think of speech as a new field. We encounter many robot customer-service systems in our daily lives, and they seem quite common already. Since you chose to work on speech, you must have identified problems that remain unresolved. What are the main challenges?

Qinyuan Cheng: Speech is indeed a field with a very long history. Some early models and companies could be considered part of the 1.0 or 2.0 era. Their main goal was to pronounce words correctly: as long as the words were understood correctly, that was enough. But beyond that, a system also needs to perceive context. For example, the same sentence may be delivered differently in different contexts. The system needs to consider the prosody before and after the sentence as well as its emotional content. The same sentence, delivered in different ways, can convey different emotions. This capability was not available in earlier systems.

Why not? Primarily because earlier approaches could not support large-model training or adopt methods that were effective for language models, such as large parameter counts, massive datasets, and post-training approaches including reinforcement learning. As a result, they could not extract these capabilities from vast amounts of data.

When speech models acquire new capabilities, they also unlock new scenarios and requirements. Consider image and video generation. In the past, people believed that many tasks could not be completed in one pass by image or video generation models. Now, however, these models can generate polished posters or illustrations for PowerPoint presentations in a single pass. Speech is similar. Once you can generate complex, high-quality speech content, new requirements will emerge alongside existing applications, such as the dubbing of short-form dramas, which many people have long wanted to create.

The Paper Technology: You mentioned that older models could only pronounce words correctly but lacked control over emotion and pauses. Can these problems be solved through training with larger parameter counts and more data?

Qinyuan Cheng: Yes. However, making effective use of massive datasets and large parameter counts also requires solving several problems. For example, the architectures of first- and second-generation models were inherently unsuitable for large-scale training. This may be because their structures were not sufficiently learnable or were not compatible with today’s advanced GPU infrastructure.

The amount of training data used across the speech field has also changed significantly. It may have started with a dozen hours, then grown to hundreds of hours and thousands of hours. Today, people commonly work with datasets on the order of millions of units, and the scale may increase further in the future, along with more precisely curated data. The reason training datasets have been able to expand continuously is that training architectures have also evolved, making it possible to train on larger amounts of data and achieve better results.

I think the solution for ultimately achieving more human-like speech or better prosody may be scaling. However, scaling requires an appropriate approach, and what we are using now may be the most suitable one.

MOSS’s Approach to Speech Models

The Paper Technology: What exactly does your approach involve?

Qinyuan Cheng: We use discrete tokens to process speech.

A typical speech generation model does not necessarily need to use a discrete representation. Some more traditional approaches directly generate spectrograms or waveforms. The problem is that these spectrograms and waveforms are not particularly easy for large models to learn. Neither their representational structure nor the infrastructure required for training is especially compatible with large-model methods. Therefore, if we want to achieve large-scale capabilities in speech through large parameter counts and massive datasets, as language models do, we need to make certain changes and conduct further research.

We have developed a relatively effective technology that converts speech into a form similar to text tokens with near-lossless quality. This allows us to take advantage of large models to generate audio that sounds more natural and more human-like.

The process of converting audio into tokens generally requires considering several factors. First is the compression rate: the tokens need to be sufficiently compact. A five-second audio clip cannot correspond to an excessive number of tokens. Second is token learnability. Ideally, the speech representation should be highly abstracted, like text, while preserving certain original characteristics of speech, such as emotion, prosody, and different timbres. We believe we have achieved a strong balance across these factors.

The Paper Technology: In your view, what is the core reason audio models are not yet as capable? Is it data, computing power, or algorithms?

Qinyuan Cheng: I think data is the most important factor. A great deal of data is still difficult to collect at scale, including labeled data, even compared with video. Video-generation requirements are currently relatively short: the longest videos people generate may be around 15 seconds. Audio, by contrast, needs to handle much longer content. Even a relatively short audio segment may last several minutes. Maintaining consistency across several minutes is difficult. You may need to concatenate multiple segments, but concatenating multiple segments introduces stability problems.

Data annotation is also very important. If you provide only an audio clip and ask someone to label the intensity of its emotional expression, there are still no particularly effective understanding models that can do this well.

Another important issue is evaluation. Without effective evaluation, we have no reliable handle for optimization, and the direction of improvement becomes uncertain. Automated evaluation has always been somewhat lacking in speech synthesis.

What exactly makes prosody good, or makes a performance expressive? People can usually tell by listening, but turning that judgment into an automated evaluation metric is difficult.

The Paper Technology: In which areas do you believe your models have an advantage?

Qinyuan Cheng: We believe our strongest capabilities are prosody and timbre similarity. Based on the general context of the text, the model can make appropriate pauses and choose where to place stress or reduce emphasis. For example, if the text appears to be a news script, the model will read it in a news-reading style. If it is a documentary script or resembles a radio drama, the model will use an appropriate style for that type of content.

The second area in which we perform particularly well is timbre similarity. If I provide the model with a recording of me speaking for around ten seconds, it can learn my voice. This is generally referred to as zero-shot or few-shot voice cloning.

To achieve this, the training material provided to the model must first be lossless. In addition, our architecture allows the timbre, prosody, and other characteristics of a reference audio clip to guide subsequent generation in a way that sounds more like a real person.

The Paper Technology: What problems remain unresolved, and what are you still working to solve?

Qinyuan Cheng: One area is the continued improvement of the model’s capabilities. In our view, taking speech generation as an example, its capabilities may form a pyramid. At the bottom are existing applications, such as call-center agents or common web-based text-to-speech tools. These applications may only require the model to pronounce words clearly and accurately. As you move upward, the requirements become more difficult. At the top may be a truly production-grade application, such as dubbing for films and television dramas. This could involve sound effects, highly expressive emotional speech, multiple speakers, and a complex interactive speech scenario.

The second major category involves the adjustments required for specific commercial use cases. In some cases, we need to make additional modifications based on the foundation model, such as for dialects.

The Paper Technology: Are dialects difficult because there is not enough data, or for some other reason?

Qinyuan Cheng: I think the fundamental issue is that there is not enough data. Very little dialect data can actually be collected online. So we may need to collaborate with organizations dedicated to dialect preservation and increase the amount of available data.

Annotation is another issue. Many dialects do not have an appropriate written representation. For example, I am from Xi’an, and I know that many Shaanxi dialect expressions are difficult to transcribe. This makes it difficult to establish a text-to-speech mapping.

In general, dialects that are more common online are relatively easier to learn. When a dialect is not well represented online, the amount of data often depends on how willing people in that region are to annotate it. For example, Jiangsu, Zhejiang, and Shanghai are economically developed regions, but people may be relatively less willing to perform annotation work.

The Paper Technology: What other challenges do you face during productization?

Qinyuan Cheng: One difference between academic research and entrepreneurship is that having a model with good performance is not enough. You also need to make more people aware of that performance. This may test our ability to operate within the open-source community and coordinate with the developer ecosystem.

In academic circles, people are generally very tolerant when a model is provided as a service. In commercial scenarios, however, users may stop using a model if it is unstable. This also tests the infrastructure capabilities of our backend, or what is now generally called a “model-as-a-service” platform. We are also recruiting experienced people to work on this.

There is another point worth adding. Sometimes users need a relatively complex capability, while a model provides only atomic capabilities. For example, our model can perform voice cloning and prosody restoration very well, and it can clone a voice across languages. If the original sentence is in English, I can translate it into Chinese while preserving almost the same timbre and a similar prosody. However, this is still a very specific capability, and users may not know what they can do with it. If you instead provide a feature that allows users to upload an entire interview video and directly convert it into an English version, that is the kind of functionality users actually need.

Therefore, we often need to combine our atomic capabilities into functions that are simpler for users and capable of delivering complex content.

AI programming is a good example. In the beginning, we used AI to assist with coding by copying a problem or writing a rough framework, pasting it into the ChatGPT chat box, and asking how to write a particular function or use a specific library. This still involved a relatively high barrier to entry, and the users were mostly researchers or engineers. Ordinary people might not even download a code editor, so they would not be able to do this. As AI programming becomes more convenient, however, more people will start using it.

The Paper Technology: Among domestic model companies, which companies would you like to be considered alongside? What kind of company would you like people to see you as?

Qinyuan Cheng: Internally, we tell people that we are the multimodal version of DeepSeek. Looking at the direction of omnimodality, Google is doing the best work. Compared with text, the gap between domestic and overseas capabilities may be even larger in multimodality. So, in a sense, we want to become a company like Google and make progress through the coordination of all modalities.

The Paper Technology: Which products has MOSS released so far?

Qinyuan Cheng: Most of the models we have released are open source. We do have some closed-source models, which are mainly focused on audio understanding and generation.

For video understanding, we want to build real-time interaction for long-form streaming video. Traditional video models respond on a turn-by-turn basis. In a streaming setup, the model can provide timely feedback based on visual changes. While generating text, it continuously receives the video stream. Once the video changes, it can immediately adjust its output. This creates a highly responsive form of interaction, and it is something existing models do not handle particularly well.

For understanding models, we released a foundation model called MOSS VL, which supports basic video-understanding and image-understanding tasks.

The Impact of AI’s Rapid Iteration on AI Professionals

The Paper Technology: Do you think the AI industry is highly competitive right now?

Qinyuan Cheng: Objectively speaking, it is very competitive. These days, I feel that doing anything else can be a bit of a waste of time. In theory, the field requires a significant investment of time and energy because training many large models is a highly continuous process. You need to consider many things: new technologies, new scenarios, and new applications. So perhaps it is not simply a matter of intense competition; the field genuinely requires this much time. We currently work roughly from 10 a.m. to 10 p.m., and some people on the algorithms side may invest even more time. It is a field that requires passion.

The Paper Technology: Compared with working in the laboratory at Fudan University, how has entrepreneurship changed your mindset?

Qinyuan Cheng: The biggest change is that I now work with many more people. The most difficult thing for me has been learning to collaborate with people who are more capable than I am. At university, I inevitably felt that I had to compete and prove whose achievements or projects were better. But within a team, you need to adapt to working with people who are much stronger than you. In some cases, you even need to organize and coordinate those people effectively. That is what is best for the overall project.

The Paper Technology: Do you have your own approach to leading a team?

Qinyuan Cheng: We are still exploring. In general, we respect a wide range of opinions and let experimental data guide our decisions. One thing I have been making sure of recently is that everyone receives feedback on the work they are doing, so they know how to iterate on their part of the process. For example, if someone is working on data but our model has not yet started training, how can that person know whether the data is good? We need to think about this and help find a solution so that the person can receive timely feedback.

The Paper Technology: How has the rapid development of AI affected people like you who work in the field?

Qinyuan Cheng: My immediate impression is that capable people have become even more capable, and people with strong energy and focus have also become more effective. They may use AI better and be better able to judge whether AI is producing reliable results. The difference is quite significant. For example, some people can let Codex work for three or four days and deliver a task, while others cannot. Overall, I think the barrier to learning new things has become lower for everyone.

I am also exploring what kinds of people use AI effectively. In general, I think people who are responsible and reliable tend to use AI better. Their deliverables are more dependable. People who do not naturally have these qualities may run into more problems when using AI.

The Paper Technology: What kind of people do you prefer to hire?

Qinyuan Cheng: We mainly hire young people. We have many master’s and doctoral interns, as well as undergraduate interns.

The first quality we value is reliability. By reliable, we mean that someone takes responsibility for their work and provides clear updates throughout the process so that others understand their current status. Many people are used to putting off a task until the end. They may feel embarrassed to speak up during the early and middle stages, which ultimately causes the task to fail. We encourage people to seek help promptly if they are blocked on something for 15 minutes. Do not let a task remain silently stuck; if it is not working, adjust quickly.

Another requirement is being relatively AI-native, meaning that the person uses AI tools effectively. This is comparatively easy to develop, and we can cultivate it through our organization.

We also hope to hire people who seek to make an impact. We are not particularly focused on credentialism or background. We care more about whether someone has independently done meaningful work, such as carefully studying some cutting-edge engineering frameworks. There are still relatively few people like this.

The Paper Technology: Are there any restrictions on academic background or major?

Qinyuan Cheng: For undergraduate students, their major is generally not a significant limitation. We take their level of commitment to AI very seriously. If they truly regard it as a career, we are willing to give them opportunities to grow.

Some of the people working on our evaluation pipelines come from linguistics or translation programs at Fudan University. Audio and video evaluation may require people who studied film or even directing. Some people who transitioned from the humanities to programming are even more impressive: they simply start writing code and do nothing related to their original field of study.

The Paper Technology: What advice do you have for young students?

Qinyuan Cheng: I think undergraduates are already very capable. Most of them now have a Code Agent at their disposal. So I have not really come up with any particularly good advice. It feels like things have changed too much. If I were to return to my first year of university now, I would not know how to plan my path either. Many of the major assignments and programming tasks that used to trouble us can now be handled by AI.

We may need new ways to plan growth, but I have not experienced that process myself. So I think learning to use AI effectively is critical. Another point is to adjust evaluation systems as early as possible toward downstream tasks. For people entering the AI industry, for example, it may be helpful to participate early in research projects or influential open-source projects.

The Paper Technology: What are your thoughts on the next five to ten years?

Qinyuan Cheng: I do not have particularly specific thoughts about myself, but I do have clear thoughts about models. We want to make the coordination among all modalities work well, tackle the different areas of multimodality one by one, and ultimately bring these capabilities into the real world.