← Search

Tongyu Lu

2 accepted papers

2025

Coarse-to-Fine Text-to-Music Latent Diffusion

ICASSP 2025accepted

We introduce DiscoDiff, a text-to-music generative model that utilizes two latent diffusion models to produce high-fidelity 44.1kHz music hierarchically. Our approach significantly enhances audio quality through a coarse-to-fine generation strategy, leveraging residual vector quantization from the D…

Cited by 0SourceScholar
2024

Vanishing-Point-Guided Video Semantic Segmentation of Driving Scenes

CVPR 2024highlight

The estimation of implicit cross-frame correspondences and the high computational cost have long been major challenges in video semantic segmentation (VSS) for driving scenes. Prior works utilize keyframes feature propagation or cross-frame attention to address these issues. By contrast we are the f…