arXiv · 2609.25546
Synthesis and editing of multi-instrument audio mixtures using scalar-quantised latents with MIDI Span conditioning
Abstract
Music creation often involves iterative refinement, changing selected musical details while retaining the rest. To support such refinement, we introduce SpanSynth-Edit, a flow-matching model for MIDI-guided synthesis and editing of multi-instrument audio mixtures using low-frame-rate scalar-quantised latents. MIDI Span encodes instrument-labelled note lifecycles as unordered event sets with continuous-valued attributes and pools each set into one conditioning vector per audio-latent frame. The model uses contextual audio for instrument-specific timbre guidance and supports editing by resynthesising the target region from revised MIDI. Experiments on single- and multi-instrument benchmarks show competitive performance and demonstrate within-frame onset control. We also discuss limitations of transcription-based note-adherence evaluation.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Sungkyun Chang, Keshav Bhandari, Simon Dixon, Emmanouil Benetos. 2026-09-22. Synthesis and editing of multi-instrument audio mixtures using scalar-quantised latents with MIDI Span conditioning. https://arxiv.org/abs/2609.25546
Cite the original work for its findings. Save a collection to share your selection of sources.