MuScriptor Large Turns Real Music Into MIDI
A new open model tackles multi-instrument transcription of real audio mixes, converting songs directly into editable MIDI.
Automatic music transcription has long been one of the harder problems in machine listening: pulling clean, note-level information out of a dense recording where several instruments overlap. MuScriptor Large, newly published on Hugging Face, takes aim squarely at that challenge, offering an open model for multi-instrument transcription that converts real audio mixes into MIDI.
The pitch is straightforward but ambitious. Rather than handling a single isolated instrument or clean studio stems, the model is built to work on real mixes — the kind of layered, messy audio that most listeners actually encounter. The output is symbolic MIDI, which means transcriptions can be edited, re-scored, or fed into downstream music tools.
Why it matters
Reliable audio-to-MIDI on full arrangements would be useful across a range of workflows:
- Musicians and educators who want to study or reproduce parts from recordings
- Producers looking to convert audio ideas into editable sequences
- Researchers building datasets and tools for music information retrieval
A few practical notes temper the enthusiasm. The model ships under a CC BY-NC 4.0 license, which restricts commercial use, so it lands most naturally in research and personal projects for now. The record also leaves parameter count and other technical specifications unstated, and the accompanying paper reference will be the place to look for benchmark details.
Still, an openly available model focused specifically on multi-instrument transcription of real recordings is a meaningful addition to the open music tooling landscape. For anyone working on the audio-to-symbolic pipeline, MuScriptor Large is worth a look.
Sources
- Visit
MuScriptor/muscriptor-large
Hugging Face
More in Music
Qwen Enters Music Generation With Qwen-Music
Alibaba's Qwen team debuts a text-to-song model that produces high-fidelity tracks complete with vocals.
Stability AI's Demon brings real-time music diffusion to local GPUs
An open-source engine generates audio on the fly at 25Hz, no cloud required.
HKUST Releases Audio-Omni, a Unified Audio Model
The new diffusion-based model handles speech, music, and general audio tasks like conversion and editing within a single, versatile framework.
0 comments
No comments yet. Be the first to weigh in.