A fully open-source pipeline is all it takes to run MOSS-VL-Realtime directly from model weights as a real-time video call in the browser.
We have open-sourced two key components in the additional support libraries of the MOSS-VL repository:
- sglang-omni — an inference backend built specifically for real-time interaction, keeping video calls with the model smooth even when multiple people connect simultaneously;
- realtime-demo — a complete out-of-the-box application with a browser interface, voice input and output, and long-conversation memory.
Together, they form a complete pipeline connecting the browser all the way to the GPU.
One Pipeline, Three Components
The complete pipeline consists of three open-source components:
- The model weights OpenMOSS-Team/MOSS-VL-Realtime-SGLANG provide a model that can be loaded directly;
- The specialized SGLang-Omni backend (MOSS-VL/third_party/sglang-omni at main · OpenMOSS/MOSS-VL) enables the model to perform efficient streaming inference on GPUs;
- Realtime Demo (MOSS-VL/third_party/realtime-demo at main · OpenMOSS/MOSS-VL) provides the user interface, voice capabilities, and memory mechanism.

Specialized SGLang-Omni Backend: Smooth Performance for Multiple Users
If the Demo is what users see, the specialized SGLang-Omni backend is the engine that makes everything run smoothly. It is an inference framework that we deeply adapted for MOSS-VL-Realtime on top of SGLang-Omni. It treats each video call as a persistent inference request and builds a complete real-time service solution around it.
To deploy it, simply configure the path to the model weights in SGLang format and start the service.
Multiple People Can Talk at Once Without Interfering
The challenge of real-time video calls is not simply responding quickly. The system must remain fast even when many people are using it at the same time. The backend allows multiple sessions to remain online simultaneously, with the video and conversation context of each session kept independent and isolated.
When the number of online users reaches the limit, newly joined sessions receive a clear “full” notification and exit gracefully, rather than slowing down everyone’s experience.
When GPU memory becomes constrained, the backend proactively evicts the most resource-intensive session to preserve normal use for the remaining sessions—much like a full restaurant asking guests who have finished their meals to give up their seats first.
The video and questions you send are always prioritized and will not get stuck simply because someone else is speaking. When the network becomes unstable, the server also provides buffering as a fallback, preventing frames from being silently dropped.
From One GPU to Multiple GPUs: Compute Scales Horizontally
Can’t the model fit on a single GPU? Split it across multiple GPUs for collaborative inference. Need to serve more users? Run multiple independent instance replicas. New sessions will automatically be assigned to the least busy replica, and only need to queue when all replicas are full.
Two GPUs, with each serving four sessions, require only one command:
两路单卡 DP 副本示例
python examples/run_moss_vl_realtime_server.py \
--model-path $MODEL_PATH \
--dp-size 2 --gpus 0,1 \
--max-running-requests 4
The generation process uses low-level acceleration by default, minimizing the cost of generating output token by token. Before the service starts, it also performs a warm-up so that the first user to connect does not have to wait through a lengthy initialization.
Long Conversations Are No Problem: Video Memory Works Like a Sliding Window
Video frames continue to accumulate throughout a conversation, creating the biggest resource bottleneck for long sessions. The backend uses a “sliding window”: by default, it keeps only the most recent frames and promptly releases older ones, while preserving the conversation text and semantics in full. The model still remembers what was discussed recently, but it no longer continuously pays the cost of retaining old frames.
How Fast Is It in Practice? Benchmark Results
We conducted system tests on a single NVIDIA H200 GPU. Compared with the reference implementation provided with the model, the specialized backend delivered the following results for a single session:
| Metric | HuggingFace | SGLang-Omni | Improvement |
|---|---|---|---|
| Time to see the first token after asking a question | 91.77 ms | 33.57 ms | 2.73× |
| Time to generate each token | 52.05 ms | 11.08 ms | 4.70× |
| Generation speed | 19.21 tokens/s | 90.26 tokens/s | 4.7× |
What happens when multiple people use it at the same time? With video continuously input at one frame per second and no rate limiting, performance remains stable from one to eight sessions, with four sessions as the default recommendation:
| Concurrent sessions | Average time to first token | Average time per token | Average generation speed per session |
|---|---|---|---|
| 1 | 512.5 ms | 16.95 ms | 59.0 tokens/s |
| 4 | 548.9 ms | 21.43 ms | 47.8 tokens/s |
| 8 | 728.1 ms | 39.03 ms | 25.6 tokens/s |
Realtime Demo: A Complete Application You Can Use in the Browser
The inference backend solves the problem of “computing quickly.” To enable users to have a real conversation, the system also needs an interface, voice capabilities, and memory. Realtime Demo provides all of these in one package. It is not just a short demo, but a complete application that can be deployed and used directly.
The services behind it orchestrate these capabilities into a pipeline: clearly hear what you say, understand what you show, organize a response, and then speak it aloud. The entire process is completed automatically.
Browser Frontend: Start a Video Call

No software installation is required. Open the webpage and start using it. You can show the model anything that is happening:
- Turn on the camera and talk to the model “face to face”;
- Share your screen and let it see what you are doing;
- Send image or video files directly.
The model’s response streams out token by token like subtitles. You can interrupt it at any time to ask a follow-up question. When you want to speak, simply press and hold. Automatic voice detection and automatic sending after you finish speaking are also supported.
There is no need to worry about an unstable network. After reconnecting, missed content is automatically sent again. If you refresh the page within 45 seconds, the conversation context remains available and you can continue chatting.
Memory Mechanism: Helping the Real-Time Assistant Remember “What We Discussed Yesterday”
In a real-time conversation, video frames and chat content continue to accumulate. If everything is retained, the model will quickly reach its capacity limit. This is the root cause of why real-time assistants become increasingly difficult to sustain during longer conversations.
The Demo uses a combined approach: a sliding window controls the size of the visual context, while text is consolidated into long-term memory and automatically retrieved when needed.
For visual content, the system keeps only the raw content from the most recent period and promptly releases older material. The model still remembers what was discussed recently, but it no longer continues paying the cost of retaining frames from the distant past.
For text, conversation content is consolidated into long-term memory. The system automatically determines whether memory should be retrieved based on each question you ask. If you casually say, “Where did I leave off last time?” the relevant content is automatically retrieved. Ordinary small talk is left undisturbed.
When the conversation becomes longer, the system automatically compresses earlier content into a summary once a threshold is reached, allowing the conversation to continue indefinitely. In addition, you can manually manage the current session at any time using two commands:
/compact— Manually compress the context into a summary and continue the conversation with a lighter context. Compression is not forgetting: as shown below, the model can still accurately answer questions about content discussed before compression.

/clear— Clear the context with one click, just as if starting a new session. If you ask about earlier content again, the model will honestly tell you that there is no earlier record, as shown below.

The reliability of this “consolidation + retrieval” mechanism can be verified directly. The complete process is shown below, corresponding to the two screenshots that follow:
- Create a memory: At the beginning of the conversation, ask the model to remember the number 011202. The model replies that it has remembered it;

- Have a normal conversation: Continue talking normally about the visual content for approximately three minutes;
- Compress the context: Execute
/compactto compress the context into a summary. At this point, the original text of that conversation is no longer in the model’s context; - Retrieve the memory: Ask, “What was the number I previously asked you to remember?” The system automatically determines that long-term memory needs to be queried, displays a “Recalled 1 memory” notification and the corresponding memory card, and the model correctly answers 011202 based on it.

Voice Input and Output: It Can Understand and Speak
Your speech is converted into text in real time by the open-source speech recognition model SenseVoice. It runs entirely on the CPU, and the subtitles refresh approximately once per second. Two audio input modes are available and can be switched at any time:
- VAD automatic detection: Starts listening when it detects that you begin speaking, keeping your hands free throughout;
- Push-to-talk (PTT): More reliable in noisy environments or when you need precise control over when to speak.
The entire module uses a pluggable design. In the future, replacing it with a faster and more accurate streaming recognition engine will require changing only one component.
Voice playback for responses supports both local and cloud-based solutions, which can be switched freely:
| Category | Available options |
|---|---|
| Local deployment | MOSS-TTS-Nano, Fun-CosyVoice3-0.5B, and MOSS-TTS-Realtime, all running locally and supporting timbre cloning from a sample audio clip |
| Cloud API | ElevenLabs and MiniMax, which can be connected by configuring an API Key |
The default configuration uses local speech synthesis. No cloud account is required to experience a complete voice conversation.
One-Click Deployment to Run the Complete Pipeline
A complete experience does not require manually assembling each component. Realtime Demo provides a one-click installer. From environment checks to starting the services, it takes only three steps:
bash bootstrap.sh --doctor-only # 环境体检
bash bootstrap.sh --with-memory --with-asr --with-tts # 安装全量组件
.venv/bin/python scripts/repro/run.py up --main-gpu 0 --memory-gpu 1 # 一键拉起
Hardware and deployment requirements:
- A complete experience is best supported by two GPUs—one for the main model and one for the memory module. If you have only one GPU, you can run video and text interaction alone;
- Speech recognition and speech synthesis use only the CPU and do not require an additional GPU;
- Linux and a relatively recent NVIDIA driver are required. The installer includes status checks and smoke tests to make it easy to confirm that everything is ready.
Start Your MOSS-VL Real-Time Journey
From one-click installation to real-time video calls in the browser, and from smooth multi-user concurrency to memory across long conversations, the complete pipeline from model to application is now fully open source.
The third_party directory also includes two supporting projects: a FlashAttention-3 acceleration module customized for MOSS-VL cross-modal attention, and the mainline SGLang, which supports offline batch inference for MOSS-VL. It has been merged into the official version and can be used directly.
We will further open external services in the future: the real-time pipeline’s API interface and online Playground. At that point, you will be able to experience real-time interaction by opening a webpage or calling an API without providing your own GPU. Stay tuned.
Related Resources
- Model weights: OpenMOSS-Team/MOSS-VL-Realtime-SGLANG
- Model repository: OpenMOSS/MOSS-VL
- Specialized SGLang-Omni backend: third_party/sglang-omni
- Realtime Demo: third_party/realtime-demo