← Search

Shuhui Wang

55 accepted papers

2026

ActiveScope: Actively Seeking and Correcting Perception for MLLMs

ICML 2026poster

Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in vision-language understanding, yet they still struggle with fine-grained perception in high-resolution images. While existing training-free methods typically rely on attention-based localization or coarse-to-fine s…

Cited by 0SourceScholar
2026

Adaptive Nonlinear Compression for Large Foundation Models

ICLR 2026poster

Despite achieving superior performance, large foundation models (LFMs) have substantial memory requirements, leading to a growing demand for model compression methods. While low-rank approximation presents a promising hardware-friendly solution, existing linear methods suffer significant information…

Cited by 0SourceScholar
2026

Adaptive Recurrent Message Passing for Test Time Computing on Graphs

ICML 2026poster

Pre-trained foundation models have demonstrated remarkable success in many domains, enabling a unified backbone to generalize across diverse downstream tasks. However, extending this paradigm to graph learning remains challenging due to the intrinsic mismatch between graph data and fixed architectur…

Cited by 0SourceScholar
2026

Global Directional Priors with Local Statistical Validation for Scalable Causal Discovery

ICML 2026poster

Constraint-based causal discovery relies on conditional independence (CI) tests whose reliability degrades as conditioning sets grow, particularly in hub-dominated graphs. Existing methods constrain adjacency or global structure, but leave conditioning-set dimensionality uncontrolled. In this paper,…

Cited by 0SourceScholar
2026

Locate-then-Sparsify: Attribution Guided Sparse Strategy for Visual Hallucination Mitigation

CVPR 2026

Despite the significant advancements in Large Vision-Language Models (LVLMs), their tendency to generate hallucinations undermines reliability and restricts broader practical deployment. Among the hallucination mitigation methods, feature steering emerges as a promising approach that reduces erroneo

Cited by 0SourcecodeScholar
2026

When LLMs Develop Languages: Symbolic Communication for Efficient Multi-Agent Reasoning

ICML 2026poster

Chain-of-Thought (CoT) prompting improves large language models (LLMs) on difficult reasoning tasks, but it generates long natural-language rationales that are poorly optimized towards higher-level machine efficiency and intelligence. We propose *Communicative Language Symbolism Routing* (CLSR), a t…

Cited by 0SourceScholar
2025

Dis²Booth: Learning Image Distribution with Disentangled Features for Text-to-Image Diffusion Models

AAAI 2025technical

Personalized image generation enables customized content creation based on the text-to-image diffusion models.However, existing personalization methods focus on fine-tuning generative models to learn to generate specific single individuals or concepts, such as an image of a specific Corgi, but are u…

Cited by 0SourcePDFScholar
2025

Divide-and-Conquer: Tree-structured Strategy with Answer Distribution Estimator for Goal-Oriented Visual Dialogue

AAAI 2025technical

Goal-oriented visual dialogue involves multi-round interaction between artificial agents, which has been of remarkable attention due to its wide applications. Given a visual scene, this task occurs when a Questioner asks an action-oriented question and an Answerer responds with the intent of letting…

2025

Edit Less, Achieve More: Dynamic Sparse Neuron Masking for Lifelong Knowledge Editing in LLMs

NeurIPS 2025poster

Lifelong knowledge editing enables continuous, precise updates to outdated knowledge in large language models (LLMs) without computationally expensive full retraining. However, existing methods often accumulate errors throughout the editing process, causing a gradual decline in both editing accuracy…

Cited by 0SourceScholar
2025

Enhancing Pre-trained Representation Classifiability can Boost its Interpretability

ICLR 2025spotlight

The visual representation of a pre-trained model prioritizes the classifiability on downstream tasks, while the widespread applications for pre-trained visual models have posed new requirements for representation interpretability. However, it remains unclear whether the pre-trained representations c…

2025

Image-to-video Adaptation with Outlier Modeling and Robust Self-learning

AAAI 2025technical

The image-to-video adaptation task seeks to effectively harness both labeled images and unlabeled videos for achieving effective video recognition. The modality gap of the image and video modalities and the domain discrepancy across the two domains are the two essential challenges in this task. Exis…

2025

Learning Fine-Grained Representations through Textual Token Disentanglement in Composed Video Retrieval

ICLR 2025poster

With the explosive growth of video data, finding videos that meet detailed requirements in large datasets has become a challenge. To address this, the composed video retrieval task has been introduced, enabling users to retrieve videos using complex queries that involve both visual and textual infor…

2025

MSR: A Multifaceted Self-Retrieval Framework for Microscopic Cascade Prediction

AAAI 2025technical

The microscopic cascade prediction task has wide applications in downstream areas like ''rumor detection''. Its goal is to forecast the diffusion routines of information cascade within networks. Existing works typically formulate it as a classification task, which fails to well align with the Social…

Cited by 0SourcePDFScholar
2025

Masked Temporal Interpolation Diffusion for Procedure Planning in Instructional Videos

ICLR 2025poster

In this paper, we address the challenge of procedure planning in instructional videos, aiming to generate coherent and task-aligned action sequences from start and end visual observations. Previous work has mainly relied on text-level supervision to bridge the gap between observed states and unobser…

Cited by 0SourcePDFScholar
2025

Relieving the Over-Aggregating Effect in Graph Transformers

NeurIPS 2025poster

Graph attention has demonstrated superior performance in graph learning tasks. However, learning from global interactions can be challenging due to the large number of nodes. In this paper, we discover a new phenomenon termed over-aggregating. Over-aggregating arises when a large volume of messages…

Cited by 0SourceScholar
2025

VL-SAE: Interpreting and Enhancing Vision-Language Alignment with a Unified Concept Set

NeurIPS 2025poster

The alignment of vision-language representations endows current Vision-Language Models (VLMs) with strong multi-modal reasoning capabilities. However, the interpretability of the alignment component remains uninvestigated due to the difficulty in mapping the semantics of multi-modal representations…

Cited by 0SourceScholar
2025

Video Language Model Pretraining with Spatio-temporal Masking

CVPR 2025poster

The development of self-supervised video-language models based on mask learning has significantly advanced downstream video tasks. These models leverage masked reconstruction to facilitate joint learning of visual and linguistic information. However, recent study reveals that reconstructing image fe…

Cited by 0SourcePDFScholar
2024

Bias-Conflict Sample Synthesis and Adversarial Removal Debias Strategy for Temporal Sentence Grounding in Video

AAAI 2024technical

Temporal Sentence Grounding in Video (TSGV) is troubled by dataset bias issue, which is caused by the uneven temporal distribution of the target moments for samples with similar semantic components in input videos or query texts. Existing methods resort to utilizing prior knowledge about bias to art…

2024

CTSM: Combining Trait and State Emotions for Empathetic Response Model

COLING 2024main

Empathetic response generation endeavors to empower dialogue systems to perceive speakers’ emotions and generate empathetic responses accordingly. Psychological research demonstrates that emotion, as an essential factor in empathy, encompasses trait emotions, which are static and context-independent…

2024

Confusing Pair Correction Based on Category Prototype for Domain Adaptation under Noisy Environments

AAAI 2024technical

In this paper, we address unsupervised domain adaptation under noisy environments, which is more challenging and practical than traditional domain adaptation. In this scenario, the model is prone to overfitting noisy labels, resulting in a more pronounced domain shift and a notable decline in the ov…

2024

Data-free Neural Representation Compression with Riemannian Neural Dynamics

ICML 2024oral

Neural models are equivalent to dynamic systems from a physics-inspired view, implying that computation on neural networks can be interpreted as the dynamical interactions between neurons. However, existing work models neuronal interaction as a weight-based linear transformation, and the nonlinearit…

Cited by 1SourcePDFScholar
2024

Expanding Sparse Tuning for Low Memory Usage

NeurIPS 2024poster

Parameter-efficient fine-tuning (PEFT) is an effective method for adapting pre-trained vision models to downstream tasks by tuning a small subset of parameters. Among PEFT methods, sparse tuning achieves superior performance by only adjusting the weights most relevant to downstream tasks, rather tha…

2024

Learning Invariant Representation with Consistency and Diversity for Semi-Supervised Source Hypothesis Transfer

ICASSP 2024accepted

Semi-supervised Domain adaptation (SSDA) has shown promising results by leveraging unlabeled data and limited labeled samples in the target domain. However, accessibility to source data is hindered by data privacy concerns, giving rise to Semi-supervised Source Hypothesis Transfer (SSHT). Integratin…

Cited by 0SourceScholar
2024

R&B: Region and Boundary Aware Zero-shot Grounded Text-to-image Generation

ICLR 2024poster

Recent text-to-image (T2I) diffusion models have achieved remarkable progress in generating high-quality images given text-prompts as input. However, these models fail to convey appropriate spatial composition specified by a layout instruction. In this work, we probe into zero-shot grounded T2I gene…

2024

Towards Dynamic Message Passing on Graphs

NeurIPS 2024poster

Message passing plays a vital role in graph neural networks (GNNs) for effective feature learning. However, the over-reliance on input topology diminishes the efficacy of message passing and restricts the ability of GNNs. Despite efforts to mitigate the reliance, existing study encounters message-pa…

2023

All in a Row: Compressed Convolution Networks for Graphs

ICML 2023poster

Compared to Euclidean convolution, existing graph convolution methods generally fail to learn diverse convolution operators under limited parameter scales and depend on additional treatments of multi-scale feature extraction. The challenges of generalizing Euclidean convolution to graphs arise from…

2023

Exploiting Completeness and Uncertainty of Pseudo Labels for Weakly Supervised Video Anomaly Detection

CVPR 2023poster

Weakly supervised video anomaly detection aims to identify abnormal events in videos using only video-level labels. Recently, two-stage self-training methods have achieved significant improvements by self-generating pseudo labels and self-refining anomaly scores with these labels. As the pseudo labe…

Cited by 92SourcePDFScholar
2023

ImageNet-E: Benchmarking Neural Network Robustness via Attribute Editing

CVPR 2023poster

Recent studies have shown that higher accuracy on ImageNet usually leads to better robustness against different corruptions. In this paper, instead of following the traditional research paradigm that investigates new out-of-distribution corruptions or perturbations deep models may encounter, we cond…

2023

The Euclidean Space is Evil: Hyperbolic Attribute Editing for Few-shot Image Generation

ICCV 2023poster

Few-shot image generation is a challenging task since it aims to generate diverse new images for an unseen category with only a few images. Existing methods suffer from the trade-off between the quality and diversity of generated images. To tackle this problem, we propose Hyperbolic Attribute Editin…

Cited by 19PDFcodeScholar
2022

Attribute Group Editing for Reliable Few-Shot Image Generation

CVPR 2022poster

Few-shot image generation is a challenging task even using the state-of-the-art Generative Adversarial Networks (GANs). Due to the unstable GAN training process and the limited training data, the generated images are often of low quality and low diversity. In this work, we propose a new "editing-bas…

Cited by 36PDFcodeScholar
2022

Hierarchical Modular Network for Video Captioning

CVPR 2022poster

Video captioning aims to generate natural language descriptions according to the content, where representation learning plays a crucial role. Existing methods are mainly developed within the supervised learning framework via word-by-word comparison of the generated caption against the ground-truth t…

Cited by 113PDFcodeScholar
2022

Learning Linguistic Association towards Efficient Text-Video Retrieval

ECCV 2022poster

"Text-video retrieval attracts growing attention recently. A dominant approach is to learn a common space for aligning two modalities. However, video deliver richer content than text in general situations and captions usually miss certain events or details in the video. The information imbalance bet…

2022

Unsupervised Coherent Video Cartoonization with Perceptual Motion Consistency

AAAI 2022technical

In recent years, creative content generations like style transfer and neural photo editing have attracted more and more attention. Among these, cartoonization of real-world scenes has promising applications in entertainment and industry. Different from image translations focusing on improving the st…

2021

Greedy Gradient Ensemble for Robust Visual Question Answering

ICCV 2021poster

Language bias is a critical issue in Visual Question Answering (VQA), where models often exploit dataset biases for the final decision without considering the image information. As a result, they suffer from performance drop on out-of-distribution data and inadequate visual explanation. Based on exp…

Cited by 78PDFcodeScholar
2021

QAIR: Practical Query-Efficient Black-Box Attacks for Image Retrieval

CVPR 2021poster

We study the query-based attack against image retrieval to evaluate its robustness against adversarial examples under the black-box setting, where the adversary only has query access to the top-k ranked unlabeled images from the database. Compared with query attacks in image classification, which pr…

Cited by 64PDFcodeScholar
2020

A Structured Latent Variable Recurrent Network With Stochastic Attention For Generating Weibo Comments

IJCAI 2020poster

Building intelligent agents to generate realistic Weibo comments is challenging. For such realistic Weibo comments, the key criterion is improving diversity while maintaining coherency. Considering that the variability of linguistic comments arises from multi-level sources, including both discourse-…

2020

Gradually Vanishing Bridge for Adversarial Domain Adaptation

CVPR 2020poster

In unsupervised domain adaptation, rich domain-specific characteristics bring great challenge to learn domain-invariant representations. However, domain discrepancy is considered to be directly minimized in existing solutions, which is difficult to achieve in practice. Some methods alleviate the dif…

Cited by 352PDFcodeScholar
2020

Interpretable Visual Reasoning via Probabilistic Formulation under Natural Supervision

ECCV 2020poster

Visual reasoning is crucial for visual question answering (VQA). However, without labelled programs, implicit reasoning under natural supervision is still quite challenging and previous models are hard to interpret. In this paper, we rethink implicit reasoning process in VQA, and propose a new formu…

2020

Label Decoupling Framework for Salient Object Detection

CVPR 2020poster

To get more accurate saliency maps, recent methods mainly focus on aggregating multi-level features from fully convolutional network (FCN) and introducing edge information as auxiliary supervision. Though remarkable progress has been achieved, we observe that the closer the pixel is to the edge, the…

Cited by 390PDFcodeScholar
2020

Parsing-Based View-Aware Embedding Network for Vehicle Re-Identification

CVPR 2020poster

Vehicle Re-Identification is to find images of the same vehicle from various views in the cross-camera scenario. The main challenges of this task are the large intra-instance distance caused by different views and the subtle inter-instance discrepancy caused by similar vehicles. In this paper, we pr…

Cited by 255PDFcodeScholar
2020

State-Relabeling Adversarial Active Learning

CVPR 2020oral

Active learning is to design label-efficient algorithms by sampling the most representative samples to be labeled by an oracle. In this paper, we propose a state relabeling adversarial active learning model (SRAAL), that leverages both the annotation and the labeled/unlabeled state information for d…

Cited by 160PDFScholar
2020

Towards Discriminability and Diversity: Batch Nuclear-Norm Maximization Under Label Insufficient Situations

CVPR 2020oral

The learning of the deep networks largely relies on the data with human-annotated labels. In some label insufficient situations, the performance degrades on the decision boundary with high data density. A common solution is to directly minimize the Shannon Entropy, but the side effect caused by entr…

Cited by 489PDFcodeScholar
2019

Adaptive Reconstruction Network for Weakly Supervised Referring Expression Grounding

ICCV 2019poster

Weakly supervised referring expression grounding aims at localizing the referential object in an image according to the linguistic query, where the mapping between the referential object and query is unknown in the training stage. To address this problem, we propose a novel end-to-end adaptive recon…

Cited by 110PDFcodeScholar
2019

Unsupervised Open Domain Recognition by Semantic Discrepancy Minimization

CVPR 2019poster

We address the unsupervised open domain recognition (UODR) problem, where categories in labeled source domain S is only a subset of those in unlabeled target domain T. The task is to correctly classify all samples in T including known and unknown categories. UODR is challenging due to the domain dis…

Cited by 38PDFcodeScholar
2018

Less is More: Picking Informative Frames for Video Captioning

ECCV 2018poster

In video captioning task, the best practice has been achieved by attention-based models which associate salient visual components with sentences in the video. However, existing study follows a common procedure which includes a frame-level appearance modeling and motion modeling on equal interval fra…

2017

A Graph Regularized Deep Neural Network for Unsupervised Image Representation Learning

CVPR 2017poster

Deep Auto-Encoder (DAE) has shown its promising power in high-level representation learning. From the perspective of manifold learning, we propose a graph regularized deep neural network (GR-DNN) to endue traditional DAEs with the ability of retaining local geometric structure. A deep-structured reg…

Cited by 37PDFcodeScholar
2015

Similarity Gaussian Process Latent Variable Model for Multi-Modal Data Analysis

ICCV 2015poster

Data from real applications involve multiple modalities representing content with the same semantics and deliver rich information from complementary aspects. However, relations among heterogeneous modalities are simply treated as observation-to-fit by existing work, and the parameterized cross-modal…

Cited by 34PDFcodeScholar