← Search

Chengen Lai

2 accepted papers

2024

Improving Vision and Language Concepts Understanding with Multimodal Counterfactual Samples

ECCV 2024poster

"Vision and Language (VL) models have achieved remarkable performance in a variety of multimodal learning tasks. The success of these models is attributed to learning a joint and aligned representation space of visual and text. However, recent popular VL models still struggle with concepts understan…

2024

Towards More Faithful Natural Language Explanation Using Multi-Level Contrastive Learning in VQA

AAAI 2024technical

Natural language explanation in visual question answer (VQA-NLE) aims to explain the decision-making process of models by generating natural language sentences to increase users' trust in the black-box systems. Existing post-hoc methods have achieved significant progress in obtaining a plausible exp…