Stability AI's Demon brings real-time music diffusion to local GPUs
An open-source engine generates audio on the fly at 25Hz, no cloud required.
Stability AI has released Demon, an open-source engine for real-time music generation built around diffusion models. Unlike batch-oriented systems that render a full track and hand it back when finished, Demon is designed to produce audio continuously, streaming sound as it computes rather than waiting for a complete render.
The headline figure is throughput: the project says it runs at 25Hz on a single local GPU. That update rate is what makes interactive use plausible, since it keeps the gap between input and output short enough to feel responsive rather than like a slow render job.
Why it matters
Most generative music tools today lean on remote servers and fixed-length outputs. A model that runs locally and in real time opens different doors:
- Latency: on-device inference avoids round trips to a cloud API.
- Control: continuous generation lends itself to live, interactive performance and experimentation.
- Access: an open-source release lets developers inspect, modify, and build on the system directly.
Demon arrives as an initial release, and key details — licensing terms, model size, and hardware requirements — remain light in the public record. Stability AI has been one of the more active companies pushing open audio and image generation, and a real-time music engine fits that pattern. For developers curious about interactive sound, the project page is the place to start.
Sources
- Visit
Show HN: Demon – open-source real-time music diffusion engine, 25Hz local GPU
Hacker News
More in Music
Qwen Enters Music Generation With Qwen-Music
Alibaba's Qwen team debuts a text-to-song model that produces high-fidelity tracks complete with vocals.
MuScriptor Large Turns Real Music Into MIDI
A new open model tackles multi-instrument transcription of real audio mixes, converting songs directly into editable MIDI.
HKUST Releases Audio-Omni, a Unified Audio Model
The new diffusion-based model handles speech, music, and general audio tasks like conversion and editing within a single, versatile framework.
0 comments
No comments yet. Be the first to weigh in.