Model Overview
For general multimodal tasks, MOSS-VL supports bilingual Chinese and English input across text, single images, multiple images, single videos, multiple videos, and interleaved image-text or video-text content. It can perform video captioning, video-based question answering and reasoning, cross-interval event association, action counting, key-information extraction from images, reasoning across multiple images, OCR, and document-layout parsing.
This moves video understanding from “watch the entire clip, then answer” to “answer while watching and revise while answering,” bringing it closer to real-time interaction.
Capabilities
-
Real-time interaction: The model answers questions while watching a continuous video stream. It can keep perceiving frames while generating a response, immediately correct or interrupt itself, and wait silently when there is not enough information. This moves from “watch first, answer later” to “answer while watching and revise while answering,” matching real online interactions.
-
Video understanding: Temporal consistency, action recognition, and causal reasoning are strengthened, allowing more accurate reconstruction of event sequences and cross-interval causal relationships in complex videos. Fine-grained action counting improves substantially in this generation, accurately counting repeated or continuous actions. Performance advances over the previous generation on benchmarks including VideoMME, MLVU, and VSI-bench.
-
Multimodal perception: Improved fine-grained object recognition and spatial reasoning enable more precise localization of small objects in high-resolution images and understanding of complex spatial relationships.
-
Key-information extraction and multi-image reasoning: Recognizes text in images through OCR and reasons about logical connections and comparisons across multiple images.
-
Document-layout parsing: Converts papers, reports, and other documents to Markdown by recognizing page structure and extracting information.
Benchmark
