• Multimodal
  • Video
  • Audio

MOSS-Video-and-Audio

A foundation model for high fidelity, synchronized video and audio generation.

Date
2026.01.29
Author
OpenMOSS Team
Organization
MOSI.AI

Introducing MOVA (MOSS Video and Audio), a foundation model designed to bridge the gap between silent video and immersive experiences. Conventional video generation often follows a cascaded "video first, audio second" pipeline. This fragmented process is computationally expensive, allows errors to accumulate, and makes precise audio-video synchronization difficult. MOVA instead enables native interaction between video and audio during generation, aligning every frame with its corresponding sound. From realistic lip movements across languages to sound effects and music that fit the scene, MOVA delivers a seamless, high fidelity audiovisual experience.

ELO RankingsMOVA: 1113.8, 95% confidence interval 1084.6–1144.3; LTX-2: 1074.1, 95% confidence interval 1045.1–1104.7; Ovi: 925.4, 95% confidence interval 895.2–952.9; WAN2.1+MMAudio: 886.9, 95% confidence interval 857.7–915.385090095010001050110011501113.8MOVA1074.1LTX-2925.4Ovi886.9WAN2.1+MMAudioModelELO

Error bars: 95% confidence intervals

Model Overview

To address the limitations of closed source systems such as Sora 2 and Veo 3, MOVA provides a fully open source framework for image-and-text-to-video-and-audio (IT2VA) and text-to-video-and-audio (T2VA) tasks. The model uses an asymmetric dual-tower architecture fused through bidirectional cross-attention. Its Mixture of Experts (MoE) design has 32B total parameters but activates only 18B during inference, balancing generation quality with deployment efficiency. Alongside the open source model weights, we provide a fine-grained dual-modal data processing pipeline and LoRA fine-tuning support, enabling researchers and creators to advance synchronized, cinematic synthesis.

MOVA model architecture

MOVA model architecture

Benchmark

ModelAudio-SpeechAV-AlignLip SyncASR Acc
IS↑DNSMOS↑DeSync↓IB-Score↑LSE-D↓LSE-C↑cpCER↓
LTX-23.0663.6350.4510.2137.2616.1090.194
Ovi3.6803.5160.5150.1907.4686.3780.439
WAN2.1 + MMAudio4.036–0.2600.317–––
MOVA-360p4.2693.7970.4750.2868.0986.2780.165
w/ dual CFG4.1693.6740.3510.3157.0047.8000.247
MOVA-720p3.9363.6710.4850.2778.0486.5930.159
w/ dual CFG3.8143.7510.3700.2977.0947.4520.195

Demos