ISMIR 2026

smorph: Playable Sound Morphing
with Diffusion Models

  • Annie Chu*1
  • Hugo Flores García*2
  • Johannes Imort2
  • Oriol Nieto2
  • Bryan Pardo1
  • Jordan Rudess3
  • Prem Seetharaman2
  • Justin Salamon2

1 Northwestern University

2 Adobe Research

3 Wizdom Music

* denotes equal contribution

SoundCanvas interface with four text or audio prompts around a two-dimensional morphing pad and RMS and pitch controls
SoundCanvas (examples below)

TL;DR: We introduce smorph, a training-free guidance framework for diffusion-based sound morphing that preserves temporal structure (e.g. RMS, pitch contours) while transforming timbral identity. It supports text and audio conditioning and enables interactive 1D and 2D morphing spaces.

full abstract

Sound morphing, generating intermediate sounds that transition from one sonic identity to another, can be a powerful tool for musical sound design. Existing diffusion-based morphing approaches entangle temporal structure and timbral identity, offering no mechanism to hold one fixed while transforming the other. We present smorph, a training-free guidance framework that preserves how a sound behaves over time while transforming what the sound is, allowing users to morph, for instance from brass to strings at a fixed pitch. We demonstrate across three morphing modes: prompt-to-prompt, audio-to-prompt, and audio-to-audio. Evaluations across diverse datasets show that smorph effectively produces smooth morph trajectories while substantially improving temporal-structure and source preservation over baselines, albeit with more conservative target-ward transformation in some settings. In an exploratory case study, musicians found smorph trajectories to be expressive and playable, suggesting structural anchoring can serve as a productive constraint for instrumental interaction.

Our Method. Figure shows N-prompt morphing with an external structural control signal. Our experiments (see paper experiments) focus on source → target morphing (N=2) with a self-derived structural control signal. We use this method to enable interactive 1D (N=2) and 2D (N=4) playable morphing spaces (see demos).
smorph method diagram
smorph method diagram

semantic slider: 1d morph

Drag the slider to hear a structure-preserving timbral morph between two text-defined endpoints. The RMS envelope of the source is held fixed while only sonic identity transforms.

source, α=0 drum beat
ctrl: rms
target, α=1 bubbles popping
spectrogram (blended)
loading audio…

soundcanvas: 4-prompt morphs

Click or drag on the pad to blend four text-prompt timbres in real time. Each 5×5 grid of sounds is pre-generated with smorph using a self-derived structural control signal; moving the crosshair performs structure-preserving timbral interpolation between the four corners.

loading audio…

soundcanvas: audio-to-prompt

One corner can be a real recording instead of a text prompt. This is audio-to-prompt morphing — editing a recording toward text-described targets while holding its temporal structure fixed. The source, a drum loop, anchors corner A, where it is initialized with SDEdit: its latent is lightly forward-noised and denoised back, so that corner reproduces the recording near-verbatim. Across the grid we vary the per-cell SDEdit start step (below) — the closer to corner A, the later denoising begins and the less noise is added, so more of the source survives; at the far corners the source latent is fully noised away and the cell is generated purely from its text prompt (tabla, cardboard box hits, 808 sub bass). Because the start step trades source preservation against transformation to the target, the grid sweeps continuously from the recording itself to pure prompt. Throughout, the source's RMS envelope is applied as a shared structural control, fixed across every denoising step, so every cell — including the pure-prompt corners — inherits the drum's onset structure. Timbre transforms toward each prompt while how the sound behaves over time stays anchored to the source.

per-cell SDEdit start step  (of 24 denoising steps)

  20  15  11   6   1    ← corner A · source recording, near-verbatim
  15  13   9   5   1
  11   9   7   3   0
   6   5   3   0   0
   1   1   0   0   0    ← far corner · pure text prompt
loading audio…