An open-source audio understanding model for real-world applications, supporting speech recognition, speaker and emotion analysis, environmental sound and music understanding, audio captioning, time-aware QA, and complex chain-of-thought reasoning.
A family of speech generation models built for production, supporting high-fidelity voice cloning, expressive conversational synthesis, environmental sound generation, and real-time streaming speech.
MOSS-TTS-Nano
An open-source, lightweight 100-million-parameter speech model built for real-time generation, combining voice cloning, multilingual synthesis, and CPU-friendly streaming inference. Ready to use on edge devices and in production.
A production-grade flagship TTS foundation model for real-world speech applications, supporting high-fidelity zero-shot voice cloning, stable long-form generation, duration control, multilingual and mixed-language synthesis, and fine-grained pronunciation control.
A production-ready model for long-form conversational speech, supporting podcast-quality generation, flexible multi-speaker switching, multilingual output, and zero-shot voice cloning from short reference clips.
Context-aware, multi-turn streaming TTS that generates low-latency incremental speech from previous text and user audio, maintaining coherent prosody, a stable voice, and natural conversational continuity.
An open-source voice design system that generates speaker voices directly from free-form text descriptions, with no reference audio required. Control character, emotion, and style, and use it as the voice layer for TTS.
A high-fidelity sound effects model for content creation, offering rich environmental sounds, broad coverage, and duration control. It reliably generates scene, action, and creature sounds from text prompts.
MOVA is a multimodal foundation model that natively generates high-fidelity video and audio in sync. With realistic lip synchronization and context-aware sound effects, it supports immersive multimedia content creation.
A model that produces accurate transcripts, speaker identities, and corresponding timestamps together, for multi-speaker settings such as meetings, interviews, podcasts, film, and television.
Mosi AI is dedicated to building Contextual Intelligence, redefining the next generation of human-AI interaction through multimodal foundation models. Here, you'll work alongside world-class engineers and researchers to push the boundaries of intelligent systems. The only limit isn't compute or data-it's your imagination for solving the next big problem.