Blog
MOSS-VL-Realtime End-to-End Real-Time Interaction Pipeline: Specialized SGLang-Omni Inference × Full-Stack Realtime Demo
Run MOSS-VL-Realtime model weights directly as real-time video calls in the browser with a fully open-source pipeline. The specialized inference backend supports persistent per-session requests, concurrent sessions, horizontal scaling, and sliding-window visual memory, delivering 90.26 tokens/s on a single GPU. The companion Demo integrates a browser interface, SenseVoice speech recognition, and long-term memory, and can be deployed with three commands.

Fine-Tuning MOSS-VL with ms-swift: A Practical Guide from Environment Setup to Doubling Performance Metrics
MOSS-VL has been merged into the ms-swift main branch. This article provides a complete workflow from environment setup to launching LoRA or Full SFT, using publicly available models and datasets without relying on internal file paths.

MOSS-VL Integrated into LlamaFactory: LoRA and Full-Parameter Fine-Tuning Ready Out of the Box
The MOSS-VL integration code has been merged into the LlamaFactory main branch. Whether you want to use LoRA to validate an idea at low cost or perform full-parameter fine-tuning for thorough training in a specific domain, you can reuse LlamaFactory’s familiar workflows for data preparation, training, checkpoint resumption, and inference.

Deploying a Native-Streaming 48 kHz MOSS-TTS Local Transformer v1.5 Speech Service on SGLang-Omni
We are releasing end-to-end serving support for MOSS-TTS-Local-Transformer-v1.5 on SGLang-Omni, a joint effort by the SGLang-Omni Team and the OpenMOSS Team.

Showing 4 of 4 articles