2024
BEAF: Observing BEfore-AFter Changes to Evaluate Hallucination in Vision-language Models
ECCV 2024poster
"Vision language models (VLMs) perceive the world through a combination of a visual encoder and a large language model (LLM). The visual encoder, pre-trained on large-scale vision-text datasets, provides zero-shot generalization to visual data, and the LLM endows its high reasoning ability to VLMs.…