SteerMusic: Enhanced Musical Consistency for Zero-shot Text-Guided and Personalized Music Editing

Authors

  • Xinlei Niu Australian National University, Canberra, Australia
  • Kin Wai Cheuk Sony AI, Tokyo, Japan
  • Jing Zhang Australian National University, Canberra, Australia
  • Naoki Murata Sony AI, Tokyo, Japan
  • Chieh-Hsin Lai Sony AI, Tokyo, Japan
  • Michele Mancusi Sony Europe B.V., Stuttgart, Germany
  • Woosung Choi Sony AI, Tokyo, Japan
  • Giorgio Fabbro Sony Europe B.V., Stuttgart, Germany
  • Wei-Hsiang Liao Sony AI, Tokyo, Japan
  • Charles Patrick Martin Australian National University, Canberra, Australia
  • Yuki Mitsufuji Sony AI, Tokyo, Japan

DOI:

https://doi.org/10.1609/aaai.v40i3.37181

Abstract

Music editing is an important step in music production, which has broad applications, including game development and film production. Most existing zero-shot text-guided editing methods rely on pretrained diffusion models by involving forward-backward diffusion processes. However, these methods often struggle to preserve the musical content. Additionally, text instructions alone usually fail to accurately describe the desired music. In this paper, we propose two music editing methods that improve the consistency between the original and edited music by leveraging score distillation. The first method, SteerMusic, is a coarse-grained zero-shot editing approach using delta denoising score. The second method, SteerMusic+, enables fine-grained personalized music editing by manipulating a concept token that represents a user-defined musical style. SteerMusic+ allows for the editing of music into user-defined musical styles that cannot be achieved by the text instructions alone. Experimental results show that our methods outperform existing approaches in preserving both music content consistency and editing fidelity. User studies further validate that our methods achieve superior music editing quality.

Published

2026-03-14

How to Cite

Niu, X., Cheuk, K. W., Zhang, J., Murata, N., Lai, C.-H., Mancusi, M., … Mitsufuji, Y. (2026). SteerMusic: Enhanced Musical Consistency for Zero-shot Text-Guided and Personalized Music Editing. Proceedings of the AAAI Conference on Artificial Intelligence, 40(3), 2000–2010. https://doi.org/10.1609/aaai.v40i3.37181

Issue

Section

AAAI Technical Track on Cognitive Modeling & Cognitive Systems