← Search

Qi Zhao*

3 accepted papers

2024

GRACE: Graph-Based Contextual Debiasing for Fair Visual Question Answering

ECCV 2024poster

"Large language models (LLMs) exhibit exceptional reasoning capabilities and have played significant roles in knowledge-based visual question-answering (VQA) systems. By conditioning on in-context examples and task-specific prompts, they comprehensively understand input questions and provide answers…

2024

GazeXplain: Learning to Predict Natural Language Explanations of Visual Scanpaths

ECCV 2024oral

"While exploring visual scenes, humans’ scanpaths are driven by their underlying attention processes. Understanding visual scanpaths is essential for various applications. Traditional scanpath models predict the where and when of gaze shifts without providing explanations, creating a gap in understa…

Cited by 4SourcePDFScholar
2024

Learning Chain of Counterfactual Thought for Bias-Robust Vision-Language Reasoning

ECCV 2024poster

"Despite the remarkable success of large vision-language models (LVLMs) on various tasks, their susceptibility to knowledge bias inherited from training data hinders their ability to generalize to new scenarios and limits their real-world applicability. To address this challenge, we propose the Coun…