Introduction
MOSS-VL is a core multimodal model family in the OpenMOSS ecosystem. Its cross-attention architecture separates visual encoding from cognitive reasoning, providing a unified foundation for offline image and video understanding. The model natively supports interleaved modalities and processes text, single or multiple images, single or multiple videos, and interleaved text, image, and video inputs in one pipeline without heavy preprocessing. Real-time interaction with continuous video streams is handled by MOSS-VL-Realtime in the same family.
The model has 11.3B parameters in total. Its 48-layer language decoder is initialized from Qwen3-8B, with a gated cross-attention layer inserted after every four layers. A 27-layer visual encoder processes images and video frames at native dynamic resolutions from 4K to 16.8 million pixels. The context length reaches 256K tokens. By injecting an absolute timestamp into each sampled frame and using XRoPE positional encoding designed for the cross-attention channel, MOSS-VL maps text and video frames into a unified three-dimensional (t, h, w) spatiotemporal coordinate system. This anchors reasoning to a precise time reference and allows the model to perceive the pace and duration of events accurately.
MOSS-VL builds its capabilities through systematic scaling across data, parameters, and context. Its four-stage pretraining and instruction tuning consumed approximately 1.37 trillion tokens, and it supports both Chinese and English. This release includes three related models: MOSS-VL-Instruct for offline conversations and downstream tasks, MOSS-VL-Base for continued pretraining and fine-tuning, and MOSS-VL-Realtime for real-time interaction. Together with two checkpoints from the previous 0408 generation, all five checkpoints are open source under Apache-2.0.

Capabilities
-
Video understanding: Absolute timestamps make the model natively compatible with variable frame rates. It supports fine-grained action localization, associations between events across time, and motion analysis such as speed and trajectory. On temporal reasoning benchmarks including Minerva (40.5), TOMATO (39.5), and VideoMME-Logical (17.1), it leads open source models of comparable size by more than 4.9 points. It scores 67.0 on EgoSchema and also covers long video benchmarks such as MLVU, LongVideoBench, and LVBench.
-
Multimodal perception: Native dynamic resolution enables fine-grained object recognition and spatial reasoning. Perception is its strongest capability area: MMBench-EN (88.1), POPE (89.4), and V*Bench (89.0) all rank first among models of comparable size, while MME-RealWorld (66.3) and BLINK (78.0) lead by 8.9 points.
-
Key information extraction from images and multi-image reasoning: The model recognizes text in images (OCR), reasons about logical and comparative relationships between images, and supports single images, multiple images, and interleaved text and image input.
-
Document layout parsing: It converts papers, reports, and other documents into Markdown, recognizing page structure and extracting information. It scores 88.9 on OmniDocBench, 4.0 points ahead of the next best model, as well as 89.6 on DocVQA, 87.8 on ChartQA, and 86.1 on OCRBench.
-
Visual grounding: It supports referring expression localization and spatiotemporal grounding. On the adversarial Ref-Adv benchmark, its score of 57.0 leads the next best model by 7.7 points.
-
Efficient inference and deployment: Visual tokens do not enter the decoding sequence, so serving latency grows more slowly as visual context increases. On a single H200 GPU, time to first token is 2.8–5.1 times faster than Qwen3-VL-8B, which shares the same language backbone. SGLang has official support, and FP8/NF4 quantized weights enable inference on a single GPU with 24 GB of VRAM. LlamaFactory and ms-swift also provide ready-to-use fine-tuning support.
Benchmark
Offline evaluation: MOSS-VL-Instruct compared with open source models of comparable size (bold indicates the best result for each benchmark; “–” means the official result was not reported):
| Capability area | Benchmark | MOSS-VL | Qwen3-VL-8B | Qwen2.5-VL-7B | LLaVA-OneVision-2-8B | Gemma-4-12B-IT |
|---|---|---|---|---|---|---|
| Multimodal perception | MMBench-EN | 88.1 | 84.8 | 83.2 | 85.8 | 82.7 |
| Multimodal perception | MMStar | 66 | 70.9 | 63.9 | 64.8 | 74.9 |
| Multimodal perception | MME-RealWorld | 66.3 | – | 57.4 | – | 46.9 |
| Multimodal perception | BLINK | 78 | 69.1 | 56.4 | 63.5 | 65.6 |
| Multimodal perception | POPE | 89.4 | – | 87.4 | – | 81.4 |
| Multimodal perception | MMMU (val) | 51.1 | 69.6 | 58.6 | – | 69.7 |
| Multimodal perception | V*Bench | 89 | 85.3 | – | 85.9 | 51.8 |
| Video understanding | VideoMME | 68.1 | 71.4 | 65.1 | 71.9 | 60.5 |
| Video understanding | MLVU (dev) | 76.8 | 78.1 | 70.2 | 76.6 | – |
| Video understanding | LongVideoBench | 65.9 | 68 | 56 | 66.9 | 58.2 |
| Video understanding | LVBench | 51.1 | 58 | 45.3 | 55.5 | 37.3 |
| Video understanding | EgoSchema | 67 | – | 65 | – | 62.2 |
| Video understanding | VSI-Bench | 62.2 | 59.4 | 28.3 | 70.9 | 25.9 |
| Video understanding | Minerva | 40.5 | – | – | – | 32.3 |
| Video understanding | TOMATO | 39.5 | 34.6 | – | – | 31.9 |
| Video understanding | VideoMME-Logical | 17.1 | 11.9 | 7.4 | – | 10.8 |
| Visual grounding | RefCOCO-REC | 84.4 | 91.6 | 90 | – | – |
| Visual grounding | Ref-Adv | 57 | 47.2 | 49.3 | – | – |
| Documents / OCR | DocVQA (val) | 89.6 | 96.1 | 95.7 | 95.2 | 80.8 |
| Documents / OCR | ChartQA | 87.8 | 89.6 | 87.3 | 85.9 | 51.2 |
| Documents / OCR | OCRBench | 86.1 | 89.6 | 86.4 | 78.2 | 76.9 |
| Documents / OCR | OmniDocBench | 88.9 | 84.9 | – | – | – |
| Reasoning | VisuLogic | 27.5 | 22.5 | 26 | – | – |
| Reasoning | ERQA | 45.8 | 45.8 | – | 43.3 | 40.8 |
Note: MOSS-VL was evaluated at 1 fps with up to 768 frames. Baseline results come from the respective official reports. See Table 5 of the technical report for the complete set of 39 benchmarks.

Inference efficiency: In measurements using one H200 GPU, SGLang deployment, and the same maximum generation length, MOSS-VL's time-to-first-token (TTFT) advantage over Qwen3-VL-8B, which shares the same language backbone, increases from 2.8× to 5.1× as visual context grows. Its end-to-end latency advantage rises from 1.9× to 4.3×. For the same video input, MOSS-VL carries roughly twice as many visual tokens because it does not perform temporal compression, yet it is never slower at any measured point.

Demos
-
Learn about the real-time interaction model MOSS-VL-Realtime.
Model downloads
| Model | Parameters | Context | Use case | Hugging Face | ModelScope |
|---|---|---|---|---|---|
| MOSS-VL-Instruct-0708 | 11B | 256K | Offline conversation, reasoning, and downstream tasks | Hugging Face | ModelScope |
| MOSS-VL-Base-0708 | 11B | 256K | Continued pretraining and fine-tuning | Hugging Face | ModelScope |
| MOSS-VL-Realtime | 11B | 256K | Real-time interaction with continuous video streams | Hugging Face | ModelScope |
| MOSS-VL-Instruct-0408 | 11B | 256K | Conversation, reasoning, and downstream tasks | Hugging Face | ModelScope |
| MOSS-VL-Base-0408 | 11B | 256K | Continued pretraining and fine-tuning | Hugging Face | ModelScope |