← Search

Jianhang Yao

3 accepted papers

2026

Beyond Attention Imbalance: Mitigating Hallucinations via Spectral Surgery

ICML 2026poster

While Large Vision-Language Models (LVLMs) achieves remarkable success, hallucinations remain a significant barrier to their reliable deployment. Recent studies primarily attribute these defects to cross-modal attention imbalances, with most solutions focusing on re-weighting visual tokens or suppre…

Cited by 0SourceScholar
2026

Plug, Play, and Fortify: A Low-Cost Module for Robust Multimodal Image Understanding Models

ICLR 2026poster

Missing modalities present a fundamental challenge in multimodal models, often causing catastrophic performance degradation. Our observations suggest that this fragility stems from an imbalanced learning process, where the model develops an implicit preference for certain modalities, leading to the…

Cited by 0SourceScholar
2026

VK-Det: Visual Knowledge Guided Prototype Learning for Open-Vocabulary Aerial Object Detection

AAAI 2026technical

To identify objects beyond predefined categories, open-vocabulary aerial object detection (OVAD) leverages the zero-shot capabilities of visual-language models (VLMs) to generalize from base to novel categories. Existing approaches typically utilize self-learning mechanisms with weak text supervisio

Cited by 0SourcePDFScholar