StepFun/Text → Speech
StepFun's StepAudio 3 Realtime targets live voice AI
The audio-language foundation model builds a listen-converse-think-act loop aimed at natural, low-latency spoken interaction.
Company
Releases
The audio-language foundation model builds a listen-converse-think-act loop aimed at natural, low-latency spoken interaction.
The new open model separates musical structure from sound, generating long-form tracks from text prompts with an explicit planning stage.
A single discrete autoregressive model handles speech, voice design, sound effects, and music generation.
The new open-source model handles both speech recognition and audio generation in a single, end-to-end architecture.