← Search

Wanying Xu

5 accepted papers

2026

Beyond Attention Imbalance: Mitigating Hallucinations via Spectral Surgery

ICML 2026poster

While Large Vision-Language Models (LVLMs) achieves remarkable success, hallucinations remain a significant barrier to their reliable deployment. Recent studies primarily attribute these defects to cross-modal attention imbalances, with most solutions focusing on re-weighting visual tokens or suppre…

Cited by 0SourceScholar
2026

Plug, Play, and Fortify: A Low-Cost Module for Robust Multimodal Image Understanding Models

ICLR 2026poster

Missing modalities present a fundamental challenge in multimodal models, often causing catastrophic performance degradation. Our observations suggest that this fragility stems from an imbalanced learning process, where the model develops an implicit preference for certain modalities, leading to the…

Cited by 0SourceScholar
2026

S3Net: Spatiotemporally Separated Sparse Network for Neuromorphic Vision Processing

AAAI 2026technical

Dynamic Vision Sensor (DVS) asynchronously records sparse events triggered by changes in pixel intensity, offering high temporal resolution and low latency. Existing frame-based methods process event data densely, violating its inherent sparsity and introducing computational redundancy. While asynch

Cited by 0SourcePDFScholar
2026

VK-Det: Visual Knowledge Guided Prototype Learning for Open-Vocabulary Aerial Object Detection

AAAI 2026technical

To identify objects beyond predefined categories, open-vocabulary aerial object detection (OVAD) leverages the zero-shot capabilities of visual-language models (VLMs) to generalize from base to novel categories. Existing approaches typically utilize self-learning mechanisms with weak text supervisio

Cited by 0SourcePDFScholar
2025

Lifting the Structural Morphing for Wide-Angle Images Rectification: Unified Content and Boundary Modeling

ICCV 2025poster

The mainstream approach for correcting distortions in wide-angle images typically involves a cascading process of rectification followed by rectangling. These tasks address distorted image content and irregular boundaries separately, using two distinct pipelines. However, this independent optimizati…