← Search

Yung-Hsuan Lai

5 accepted papers

2025

UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video Parsing

CVPR 2025poster

Audio-Visual Video Parsing (AVVP) entails the challenging task of localizing both uni-modal events (i.e., those occurring exclusively in either the visual or acoustic modality of a video) and multi-modal events (i.e., those occurring in both modalities concurrently). Moreover, the prohibitive cost o…

2024

RAPPER: Reinforced Rationale-Prompted Paradigm for Natural Language Explanation in Visual Question Answering

ICLR 2024poster

Natural Language Explanation (NLE) in vision and language tasks aims to provide human-understandable explanations for the associated decision-making process. In practice, one might encounter explanations which lack informativeness or contradict visual-grounded facts, known as implausibility and hall…

Cited by 2SourcePDFScholar
2024

Receler: Reliable Concept Erasing of Text-to-Image Diffusion Models via Lightweight Erasers

ECCV 2024poster

"Concept erasure in text-to-image diffusion models aims to disable pre-trained diffusion models from generating images related to a target concept. To perform reliable concept erasure, the properties of robustness and locality are desirable. The former refrains the model from producing images associ…

2024

Select and Distill: Selective Dual-Teacher Knowledge Transfer for Continual Learning on Vision-Language Models

ECCV 2024poster

"Large-scale vision-language models (VLMs) have shown a strong zero-shot generalization capability on unseen-domain data. However, adapting pre-trained VLMs to a sequence of downstream tasks often leads to the forgetting of previously learned knowledge and a reduction in zero-shot classification per…

Cited by 11SourcePDFScholar
2023

Modality-Independent Teachers Meet Weakly-Supervised Audio-Visual Event Parser

NeurIPS 2023poster

Audio-visual learning has been a major pillar of multi-modal machine learning, where the community mostly focused on its $\textit{modality-aligned}$ setting, $\textit{i.e.}$, the audio and visual modality are $\textit{both}$ assumed to signal the prediction target. With the Look, Listen, and Parse d…