MOSS-SoundEffect goes beyond generating high quality audio, bringing systematic improvements to the richness of ambient sound, the range of sound types, and duration control. Trained on more than 10,000 hours of high quality data, it reliably generates natural soundscapes, urban ambience, animal sounds, human actions, and music clips from text prompts. It serves content creation, games, film and television, data synthesis, and other use cases.
The model shares the same architecture as MOSS-TTS.