← Search

Guanzhen Li

3 accepted papers

2024

MVP-Bench: Can Large Vision-Language Models Conduct Multi-level Visual Perception Like Humans?

EMNLP 2024finding

Humans perform visual perception at multiple levels, including low-level object recognition and high-level semantic interpretation such as behavior understanding. Subtle differences in low-level details can lead to substantial changes in high-level perception. For example, substituting the shopping…

2024

V-DPO: Mitigating Hallucination in Large Vision Language Models via Vision-Guided Direct Preference Optimization

EMNLP 2024finding

Large vision-language models (LVLMs) suffer from hallucination, resulting in misalignment between the output textual response and the input visual content. Recent research indicates that the over-reliance on the Large Language Model (LLM) backbone, as one cause of the LVLM hallucination, inherently…

2023

ECHo: A Visio-Linguistic Dataset for Event Causality Inference via Human-Centric Reasoning

EMNLP 2023long findings

We introduce ECHo (Event Causality Inference via Human-Centric Reasoning), a diagnostic dataset of event causality inference grounded in visio-linguistic social scenarios. ECHo employs real-world human-centric deductive information building on a television crime drama. ECHo requires the Theory-of-Mi…

Cited by 0SourcecodeScholar