← Search

Kun Su

12 accepted papers

2026

Efficient, Property-Aligned Fan-Out Retrieval via RL-Amortized Diffusion

ICML 2026poster

Many modern retrieval problems are \emph{set-valued}: given a broad intent, the system must return a \emph{collection} of results that optimizes higher-order properties (e.g., diversity, coverage, complementarity, coherence) while staying grounded to a fixed database. Set-valued objectives are inher…

Cited by 0SourceScholar
2025

Diff4Steer: Steerable Diffusion Prior for Generative Music Retrieval with Semantic Guidance

ICASSP 2025accepted

Modern music retrieval systems often rely on fixed representations of user preferences, limiting their ability to capture users’ diverse and uncertain retrieval needs. To address this limitation, we introduce Diff4Steer, a novel generative retrieval framework that employs lightweight diffusion model…

Cited by 0SourceScholar
2025

UniMuMo: Unified Text, Music, and Motion Generation

AAAI 2025technical

We introduce UniMuMo, a unified multimodal model capable of taking arbitrary text, music, and motion data as input conditions to generate outputs across all three modalities. To address the lack of time-synchronized data, we align unpaired music and motion data based on rhythmic patterns to leverage…

2024

From Vision to Audio and Beyond: A Unified Model for Audio-Visual Representation and Generation

ICML 2024poster

Video encompasses both visual and auditory data, creating a perceptually rich experience where these two modalities complement each other. As such, videos are a valuable type of media for the investigation of the interplay between audio and visual elements. Previous studies of audio-visual modalitie…

2024

Tell What You Hear From What You See - Video to Audio Generation Through Text

NeurIPS 2024poster

The content of visual and audio scenes is multi-faceted such that a video stream can be paired with various audio streams and vice-versa. Thereby, in video-to-audio generation task, it is imperative to introduce steering approaches for controlling the generated audio. While Video-to-Audio generation…

2024

V2Meow: Meowing to the Visual Beat via Video-to-Music Generation

AAAI 2024technical

Video-to-music generation demands both a temporally localized high-quality listening experience and globally aligned video-acoustic signatures. While recent music generation models excel at the former through advanced audio codecs, the exploration of video-acoustic signatures has been confined to sp…

Cited by 12SourcePDFScholar
2023

Be Everywhere - Hear Everything (BEE): Audio Scene Reconstruction by Sparse Audio-Visual Samples

ICCV 2023poster

Fully immersive and interactive audio-visual scenes are dynamic such that the listeners and the sound emitters move and interact with each other. Reconstruction of an immersive sound experience, as it happens in the scene, requires detailed reconstruction of the audio perceived by the listener at an…

Cited by 9PDFScholar
2023

Physics-Driven Diffusion Models for Impact Sound Synthesis From Videos

CVPR 2023poster

Modeling sounds emitted from physical object interactions is critical for immersive perceptual experiences in real and virtual worlds. Traditional methods of impact sound synthesis use physics simulation to obtain a set of physics parameters that could represent and synthesize the sound. However, th…

Cited by 30SourcePDFScholar
2021

How Does it Sound?

NeurIPS 2021poster

One of the primary purposes of video is to capture people and their unique activities. It is often the case that the experience of watching the video can be enhanced by adding a musical soundtrack that is in-sync with the rhythmic features of these activities. How would this soundtrack sound? Such a…

Cited by 42SourcePDFScholar