← Search

Zehua Chen

10 accepted papers

2026

GuidedBridge: Training-freely Improving Bridge Models with Prior Guidance

ICML 2026poster

Guidance methods, e.g., classifier-free guidance (CFG) and auto-guidance (AG), have distinctively improved noise-to-data diffusion generation results. Recently, bridge models have been proposed, which present a data-to-data sampling process to exploit instructive information from clean prior represe…

Cited by 0SourceScholar
2026

Omni2Sound: Towards Unified Video-Text-to-Audio Generation

CVPR 2026

Training a unified model integrating video-to-audio (V2A), text-to-audio (T2A), and joint video-text-to-audio (VT2A) generation offers significant application flexibility, yet faces two unexplored foundational challenges: (1) the scarcity of high-quality audio captions with tight V-A-T alignment, le

Cited by 0SourceScholar
2025

DiffGAP: A Lightweight Diffusion Module in Contrastive Space for Bridging Cross-Model Gap

ICASSP 2025accepted

Recent works in cross-modal understanding and generation, notably through models like CLAP (Contrastive Language-Audio Pretraining) and CAVP (Contrastive Audio-Visual Pretraining), have significantly enhanced the alignment of text, video, and audio embeddings via a single contrastive loss. However,…

Cited by 0SourceScholar
2025

FrameBridge: Improving Image-to-Video Generation with Bridge Models

ICML 2025poster

Diffusion models have achieved remarkable progress on image-to-video (I2V) generation, while their noise-to-data generation process is inherently mismatched with this task, which may lead to suboptimal synthesis quality. In this work, we present FrameBridge. By modeling the frame-to-frames generatio…

2025

RespDiff: An End-to-End Multi-scale RNN Diffusion Model for Respiratory Waveform Estimation from PPG Signals

ICASSP 2025accepted

Respiratory rate (RR) is a critical health indicator often monitored under inconvenient scenarios, limiting its practicality for continuous monitoring. Photoplethysmography (PPG) sensors, increasingly integrated into wearable devices, offer a chance to continuously estimate RR in a portable manner.…

Cited by 0SourceScholar
2023

AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

ICML 2023poster

Text-to-audio (TTA) systems have recently gained attention for their ability to synthesize general audio based on text descriptions. However, previous studies in TTA have limited generation quality with high computational costs. In this study, we propose AudioLDM, a TTA system that is built on a lat…

2022

BinauralGrad: A Two-Stage Conditional Diffusion Probabilistic Model for Binaural Audio Synthesis

NeurIPS 2022accept

Binaural audio plays a significant role in constructing immersive augmented and virtual realities. As it is expensive to record binaural audio from the real world, synthesizing them from mono audio has attracted increasing attention. This synthesis process involves not only the basic physical warpin…

2022

Infergrad: Improving Diffusion Models for Vocoder by Considering Inference in Training

ICASSP 2022accepted

Denoising diffusion probabilistic models (diffusion models for short) require a large number of iterations in inference to achieve the generation quality that matches or surpasses the state-of-the-art generative models, which invariably results in slow inference speed. Previous approaches aim to opt…

Cited by 0SourceScholar