← Search

Siqi Lu

5 accepted papers

2026

Beyond Attention Imbalance: Mitigating Hallucinations via Spectral Surgery

ICML 2026poster

While Large Vision-Language Models (LVLMs) achieves remarkable success, hallucinations remain a significant barrier to their reliable deployment. Recent studies primarily attribute these defects to cross-modal attention imbalances, with most solutions focusing on re-weighting visual tokens or suppre…

Cited by 0SourceScholar
2026

Plug, Play, and Fortify: A Low-Cost Module for Robust Multimodal Image Understanding Models

ICLR 2026poster

Missing modalities present a fundamental challenge in multimodal models, often causing catastrophic performance degradation. Our observations suggest that this fragility stems from an imbalanced learning process, where the model develops an implicit preference for certain modalities, leading to the…

Cited by 0SourceScholar
2026

VK-Det: Visual Knowledge Guided Prototype Learning for Open-Vocabulary Aerial Object Detection

AAAI 2026technical

To identify objects beyond predefined categories, open-vocabulary aerial object detection (OVAD) leverages the zero-shot capabilities of visual-language models (VLMs) to generalize from base to novel categories. Existing approaches typically utilize self-learning mechanisms with weak text supervisio

Cited by 0SourcePDFScholar
2025

ASIGN: An Anatomy-aware Spatial Imputation Graphic Network for 3D Spatial Transcriptomics

CVPR 2025poster

Spatial transcriptomics (ST) is an emerging technology that enables medical computer vision scientists to automatically interpret the molecular profiles underlying morphological features. Currently, however, most deep learning-based ST analyses are limited to two-dimensional (2D) sections, which can…

2025

Lifting the Structural Morphing for Wide-Angle Images Rectification: Unified Content and Boundary Modeling

ICCV 2025poster

The mainstream approach for correcting distortions in wide-angle images typically involves a cascading process of rectification followed by rectangling. These tasks address distorted image content and irregular boundaries separately, using two distinct pipelines. However, this independent optimizati…