← Search

Sitong Yan

3 accepted papers

2024

Improving Vision and Language Concepts Understanding with Multimodal Counterfactual Samples

ECCV 2024poster

"Vision and Language (VL) models have achieved remarkable performance in a variety of multimodal learning tasks. The success of these models is attributed to learning a joint and aligned representation space of visual and text. However, recent popular VL models still struggle with concepts understan…

2024

Towards More Faithful Natural Language Explanation Using Multi-Level Contrastive Learning in VQA

AAAI 2024technical

Natural language explanation in visual question answer (VQA-NLE) aims to explain the decision-making process of models by generating natural language sentences to increase users' trust in the black-box systems. Existing post-hoc methods have achieved significant progress in obtaining a plausible exp…

2023

TITAN : Task-oriented Dialogues with Mixed-Initiative Interactions

IJCAI 2023poster

In multi-domain task-oriented dialogue systems, users proactively propose a series of domain-specific requests that can often be under-or over-specified, sometimes with ambiguous and cross-domain demands. System-sided initiative would be necessary to identify certain situations and appropriately int…