← Search

Manh Luong

3 accepted papers

2025

Unbiased Sliced Wasserstein Kernels for High-Quality Audio Captioning

NeurIPS 2025poster

Audio captioning systems face a fundamental challenge: teacher-forcing training creates exposure bias that leads to caption degeneration during inference. While contrastive methods have been proposed as solutions, they typically fail to capture the crucial temporal relationships between acoustic and…

Cited by 0SourceScholar
2024

Revisiting Deep Audio-Text Retrieval Through the Lens of Transportation

ICLR 2024poster

The Learning-to-match (LTM) framework proves to be an effective inverse optimal transport approach for learning the underlying ground metric between two sources of data, facilitating subsequent matching. However, the conventional LTM framework faces scalability challenges, necessitating the use of t…