StepFun's StepAudio 3 Music Plans Before It Plays
The new open model separates musical structure from sound, generating long-form tracks from text prompts with an explicit planning stage.
Category · audio
Open models that generate music and audio from text or melody — instrumentals, sound design, and full tracks, with weights you can download and tune.
10 releases
The new open model separates musical structure from sound, generating long-form tracks from text prompts with an explicit planning stage.
A single discrete autoregressive model handles speech, voice design, sound effects, and music generation.
The 3-billion-parameter model adds symbolic planning and agentic editing to open-source music generation.
The 3-billion-parameter model pairs symbolic planning with agentic editing to generate and refine full songs in English and Chinese.
The Chinese AI lab brings its music generation system to Hugging Face, expanding the roster of open audio models.
Alibaba's Qwen team debuts a text-to-song model that produces high-fidelity tracks complete with vocals.
A new open model tackles multi-instrument transcription of real audio mixes, converting songs directly into editable MIDI.
An open-source engine generates audio on the fly at 25Hz, no cloud required.
The new diffusion-based model handles speech, music, and general audio tasks like conversion and editing within a single, versatile framework.
The new model, SoulX-Singer, can replicate a singing voice from a short audio sample and supports both English and Chinese under a permissive license.