← Search

Yongbin Zheng

6 accepted papers

2026

Beyond Attention Imbalance: Mitigating Hallucinations via Spectral Surgery

ICML 2026poster

While Large Vision-Language Models (LVLMs) achieves remarkable success, hallucinations remain a significant barrier to their reliable deployment. Recent studies primarily attribute these defects to cross-modal attention imbalances, with most solutions focusing on re-weighting visual tokens or suppre…

Cited by 0SourceScholar
2026

Plug, Play, and Fortify: A Low-Cost Module for Robust Multimodal Image Understanding Models

ICLR 2026poster

Missing modalities present a fundamental challenge in multimodal models, often causing catastrophic performance degradation. Our observations suggest that this fragility stems from an imbalanced learning process, where the model develops an implicit preference for certain modalities, leading to the…

Cited by 0SourceScholar
2026

VK-Det: Visual Knowledge Guided Prototype Learning for Open-Vocabulary Aerial Object Detection

AAAI 2026technical

To identify objects beyond predefined categories, open-vocabulary aerial object detection (OVAD) leverages the zero-shot capabilities of visual-language models (VLMs) to generalize from base to novel categories. Existing approaches typically utilize self-learning mechanisms with weak text supervisio

Cited by 0SourcePDFScholar
2025

Lifting the Structural Morphing for Wide-Angle Images Rectification: Unified Content and Boundary Modeling

ICCV 2025poster

The mainstream approach for correcting distortions in wide-angle images typically involves a cascading process of rectification followed by rectangling. These tasks address distorted image content and irregular boundaries separately, using two distinct pipelines. However, this independent optimizati…

2024

Efficient-PIP: Large-scale Pixel-level Aligned Image Pair Generation for Cross-time Infrared-RGB Translation

IROS 2024poster

Generative models are gaining momentum in both academic and industrial applications driven by the availability of large-scale datasets, especially in tasks involving Image-to-Image Translation. Meanwhile, poor human perception of nighttime environment has led to a demand for translation from night-v…

Cited by 0SourcecodeScholar
2024

Weakly-Supervised Depth Completion during Robotic Micromanipulation from a Monocular Microscopic Image

ICRA 2024poster

Obtaining three-dimensional information, especially the z-axis depth information, is crucial for robotic micromanipulation. Due to the unavailability of depth sensors such as lidars in micromanipulation setups, traditional depth acquisition methods such as depth from focus or depth from defocus dire…

Cited by 0SourceScholar