← Search

Wei Suo

10 accepted papers

2026

Beyond Attention Imbalance: Mitigating Hallucinations via Spectral Surgery

ICML 2026poster

While Large Vision-Language Models (LVLMs) achieves remarkable success, hallucinations remain a significant barrier to their reliable deployment. Recent studies primarily attribute these defects to cross-modal attention imbalances, with most solutions focusing on re-weighting visual tokens or suppre…

Cited by 0SourceScholar
2026

Hallucination-aware Intermediate Representation Editing in Large Vision-Lanugage Models

ICLR 2026poster

Large Vision-Language Models have demonstrated exceptional performance in multimodal reasoning and complex scene understanding. However, these models still face significant hallucination issues, where outputs contradict visual facts. Recent research on hallucination mitigation has focused on retrain…

Cited by 0SourcecodeScholar
2026

Understanding and Mitigating Hallucinations in Multimodal Chain-of-Thought Models

CVPR 2026

Multimodal Chain-of-Thought (MCoT) models have demonstrated impressive capability in complex visual reasoning tasks. Unfortunately, recent studies reveal that they suffer from severe hallucination problems due to diminished visual attention during the generation process.However, visual attention dec

Cited by 0SourcecodeScholar
2025

Octopus: Alleviating Hallucination via Dynamic Contrastive Decoding

CVPR 2025highlight

Large Vision-Language Models (LVLMs) have obtained impressive performance in visual content understanding and multi-modal reasoning. Unfortunately, these large models suffer from serious hallucination problems and tend to generate fabricated responses. Recently, several Contrastive Decoding (CD) str…

2025

Pruning All-Rounder: Rethinking and Improving Inference Efficiency for Large Vision Language Models

ICCV 2025poster

Although Large Vision-Language Models (LVLMs) have achieved impressive results, their high computational costs pose a significant barrier to wide application. To enhance inference efficiency, most existing approaches can be categorized as parameter-dependent or token-dependent strategies to reduce c…

2024

C3L: Content Correlated Vision-Language Instruction Tuning Data Generation via Contrastive Learning

IJCAI 2024poster

Vision-Language Instruction Tuning (VLIT) is a critical training phase for Large Vision-Language Models (LVLMs). With the improving capabilities of open-source LVLMs, researchers have increasingly turned to generate VLIT data by using open-source LVLMs and achieved significant progress. However, suc…

Cited by 0SourcePDFScholar
2024

Rethinking and Improving Visual Prompt Selection for In-Context Learning Segmentation Framework

ECCV 2024poster

"As a fundamental and extensively studied task in computer vision, image segmentation aims to locate and identify different semantic concepts at the pixel level. Recently, inspired by In-Context Learning (ICL), several generalist segmentation frameworks have been proposed, providing a promising para…

2023

S3C: Semi-Supervised VQA Natural Language Explanation via Self-Critical Learning

CVPR 2023poster

VQA Natural Language Explanation (VQA-NLE) task aims to explain the decision-making process of VQA models in natural language. Unlike traditional attention or gradient analysis, free-text rationales can be easier to understand and gain users' trust. Existing methods mostly use post-hoc or self-ratio…

Cited by 10SourcePDFScholar
2022

A Simple and Robust Correlation Filtering Method for Text-Based Person Search

ECCV 2022poster

"Text-based person search aims to associate pedestrian images with natural language descriptions. In this task, extracting differentiated representations and aligning them among identities and descriptions is an essential yet challenging problem. Most of the previous methods depend on additional lan…

2021

Proposal-free One-stage Referring Expression via Grid-Word Cross-Attention

IJCAI 2021poster

Referring Expression Comprehension (REC) has become one of the most important tasks in visual reasoning, since it is an essential step for many vision-and-language tasks such as visual question answering. However, it has not been widely used in many downstream tasks because it suffers 1) two-stage m…

Cited by 12SourcePDFScholar