← Search

Haigang Zhang

2 accepted papers

2025

Seeing More, Saying More: Lightweight Language Experts are Dynamic Video Token Compressors

EMNLP 2025

Recent advancements in large video-language models have revolutionized video understanding tasks. However, their efficiency is significantly constrained by processing high volumes of visual tokens. Existing token compression strategies apply a fixed compression ratio, ignoring the variability in sem

2024

Reflective Instruction Tuning: Mitigating Hallucinations in Large Vision-Language Models

ECCV 2024poster

"Large vision-language models (LVLMs) have shown promising performance on a variety of vision-language tasks. However, they remain susceptible to hallucinations, generating outputs misaligned with visual content or instructions. While various mitigation strategies have been proposed, they often negl…

Cited by 7SourcePDFScholar