ICASSP 2025accepted0 citations

Audio Texture Manipulation by Exemplar-Based Analogy

Kan Jen Cheng, Tingle Li, Gopala Anumanchipalli

Abstract

Audio texture manipulation involves modifying the perceptual characteristics of a sound to achieve specific transformations, such as adding, removing, or replacing auditory elements. In this paper, we propose an exemplar-based analogy model for audio texture manipulation. Instead of conditioning on text-based instructions, our method uses paired speech examples, where one clip represents the original sound and another illustrates the desired transformation. The model learns to apply the same transformation to new input, allowing for the manipulation of sound textures. We construct a quadruplet dataset representing various editing tasks, and train a latent diffusion model in a self-supervised manner. We show through quantitative evaluations and perceptual studies that our model outperforms text-conditioned baselines and generalizes to real-world, out-of-distribution, and non-speech scenarios.

BibTeX
@inproceedings{icassp2025_audiotexturemani,
  title = {Audio Texture Manipulation by Exemplar-Based Analogy},
  author = {Kan Jen Cheng and Tingle Li and Gopala Anumanchipalli},
  booktitle = {ICASSP 2025},
  year = {2025}
}