← Search

Sungkyung Kim

2 accepted papers

2024

Exploiting Semantic Reconstruction to Mitigate Hallucinations in Vision-Language Models

ECCV 2024poster

"Hallucinations in vision-language models pose a significant challenge to their reliability, particularly in the generation of long captions. Current methods fall short of accurately identifying and mitigating these hallucinations. To address this issue, we introduce ESREAL, a novel unsupervised rei…

Cited by 5SourcePDFScholar
2024

Towards Efficient Visual-Language Alignment of the Q-Former for Visual Reasoning Tasks

EMNLP 2024finding

Recent advancements in large language models have demonstrated enhanced capabilities in visual reasoning tasks by employing additional encoders for aligning different modalities. While the Q-Former has been widely used as a general encoder for aligning several modalities including image, video, audi…