Abstract
Editing complex real-world sound scenes is difficult because individual soundsources overlap in time. Generative models can fill-in missing or corrupteddetails based on their strong prior understanding of the data domain. Wepresent a system for editing individual sound events within complex scenes ableto delete, insert, and enhance individual sound events based on textual editdescriptions (e.g., ``enhance Door'') and a graphical representation of theevent timing derived from an ``event roll'' transcription. We present anencoder-decoder transformer working on SoundStream representations, trained onsynthetic (input, desired output) audio example pairs formed by adding isolatedsound events to dense, real-world backgrounds. Evaluation reveals theimportance of each part of the edit descriptions -- action, class, timing. Ourwork demonstrates ``recomposition'' is an important and practical application.