← Search

Qinghua Hu

76 accepted papers

2026

ALSO: Adversarial Online Strategy Optimization for Social Agents

ICML 2026poster

Social simulation provides a compelling testbed for studying social intelligence, where agents interact through multi-turn dialogues under evolving contexts and strategically adapting opponents. Such environments are inherently non-stationary, requiring agents to dynamically adjust their strategies …

Cited by 0SourceScholar
2026

CtrlFuse: Mask-Prompt Guided Controllable Infrared and Visible Image Fusion

AAAI 2026technical

Infrared and visible image fusion generates all-weather perception-capable images by combining complementary modalities, enhancing environmental awareness for intelligent unmanned systems. Existing methods either focus on pixel-level fusion while overlooking downstream task adaptability or implicitl

Cited by 0SourcePDFScholar
2026

Dream-IF: Dynamic Relative EnhAnceMent for Image Fusion

AAAI 2026technical

Image fusion aims to integrate comprehensive information from images acquired through multiple sources. However, images captured by diverse sensors often encounter various degradations that can negatively affect fusion quality. Traditional fusion methods generally treat image enhancement and fusion

Cited by 0SourcePDFScholar
2026

Less Is More: Rethinking Parameter-Efficient Fine-Tuning from a Subtractive Perspective

AAAI 2026technical

Currently, pretrained models are rapidly scaling in size, which substantially increases the cost of fine-tuning them for downstream tasks. To address this challenge, parameter-efficient fine-tuning (PEFT) methods have been developed to optimize a minimal set of parameters for adaptation. While curre

Cited by 0SourcePDFScholar
2026

Test-Time Multi-Prompt Adaptation for Open-Vocabulary Remote Sensing Image Segmentation

CVPR 2026

The rise of vision-language models (VLMs) has driven the initial exploration of open-vocabulary remote sensing image semantic segmentation (OVRSIS), enabling recognition of unseen categories in complex Earth observation scenes. However, existing methods primarily focus on enhancing visual representa

Cited by 0SourcecodeScholar
2025

Asymmetric Factorized Bilinear Operation for Vision Transformer

ICLR 2025poster

As a core component of Transformer-like deep architectures, a feed-forward network (FFN) for channel mixing is responsible for learning features of each token. Recent works show channel mixing can be enhanced by increasing computational burden or can be slimmed at the sacrifice of performance. Altho…

Cited by 0SourcePDFScholar
2025

Asymmetric Reinforcing Against Multi-Modal Representation Bias

AAAI 2025technical

The strength of multimodal learning lies in its ability to integrate information from various sources, providing rich and comprehensive insights. However, in real-world scenarios, multi-modal systems often face the challenge of dynamic modality contributions, the dominance of different modalities ma…

2025

CoE: Chain-of-Explanation via Automatic Visual Concept Circuit Description and Polysemanticity Quantification

CVPR 2025poster

Explainability is a critical factor influencing the wide deployment of deep vision models (DVMs). Concept-based post-hoc explanation methods can provide both global and local insights into model decisions. However, current methods in this field face challenges in that they are inflexible to automati…

2025

DALIP: Distribution Alignment-based Language-Image Pre-Training for Domain-Specific Data

ICCV 2025poster

Recently, Contrastive Language-Image Pre-training (CLIP) has shown promising performance in domain-specific data (e.g., biology), and has attracted increasing research attention. Existing works generally focus on collecting extensive domain-specific data and directly tuning the original CLIP models.…

2025

Decoupled Multi-Predictor Optimization for Inference-Efficient Model Tuning

ICCV 2025poster

Recently, remarkable progress has been made in large-scale pre-trained model tuning, and inference efficiency is becoming more crucial for practical deployment. Early exiting in conjunction with multi-stage predictors, when cooperated with a parameter-efficient fine-tuning strategy, offers a straigh…

2025

Dynamic Personality in LLM Agents: A Framework for Evolutionary Modeling and Behavioral Analysis in the Prisoner’s Dilemma

ACL 2025finding

Using Large Language Model agents to simulate human game behaviors offers valuable insights for human social psychology in anthropomorphic AI research. While current models rely on static personality traits, real-world evidence shows personality evolves through environmental feedback. Recent work in…

Cited by 0SourcePDFScholar
2025

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark

ICLR 2025poster

The dynamic imbalance of the fore-background is a major challenge in video object counting, which is usually caused by the sparsity of target objects. This remains understudied in existing works and often leads to severe under-/over-prediction errors. To tackle this issue in video object counting, w…

Cited by 1SourcePDFScholar
2025

Graphs Help Graphs: Multi-Agent Graph Socialized Learning

NeurIPS 2025poster

Graphs in the real world are fragmented and dynamic, lacking collaboration akin to that observed in human societies. Existing paradigms present collaborative information collapse and forgetting, making collaborative relationships poorly autonomous and interactive information insufficient. Moreover,…

Cited by 0SourcecodeScholar
2025

Hierarchical Classification Auxiliary Network for Time Series Forecasting

AAAI 2025technical

Deep learning has significantly advanced time series forecasting through its powerful capacity to capture sequence relationships. However, training these models with the Mean Square Error (MSE) loss often results in over-smooth predictions, making it challenging to handle the complexity and learn hi…

2025

Learning Pattern-Specific Experts for Time Series Forecasting Under Patch-level Distribution Shift

NeurIPS 2025poster

Time series forecasting, which aims to predict future values based on historical data, has garnered significant attention due to its broad range of applications. However, real-world time series often exhibit heterogeneous pattern evolution across segments, such as seasonal variations, regime changes…

Cited by 0SourcecodeScholar
2025

Online Clustering of Dueling Bandits

ICML 2025poster

The contextual multi-armed bandit (MAB) is a widely used framework for problems requiring sequential decision-making under uncertainty, such as recommendation systems. In applications involving a large number of users, the performance of contextual MAB can be significantly improved by facilitating c…

Cited by 0SourcePDFScholar
2025

Patch-wise Structural Loss for Time Series Forecasting

ICML 2025poster

Time-series forecasting has gained significant attention in machine learning due to its crucial role in various domains. However, most existing forecasting models rely heavily on point-wise loss functions like Mean Squared Error, which treat each time step independently and neglect the structural de…

2025

Public Opinion Field Effect and Hawkes Process Join Hands for Information Popularity Prediction

AAAI 2025technical

Information popularity prediction, aiming to predict the growth of user participation in a trending topic diffusion, is a fundamental task in social networks. Existing methods often treat information diffusion as a single independent process, ignoring the ``public opinion field effect'' where multip…

2025

Reducing Class-wise Confusion for Incremental Learning with Disentangled Manifolds

CVPR 2025poster

Class incremental learning (CIL) aims to enable models to continuously learn new classes without catastrophically forgetting old ones. A promising direction is to learn and use prototypes of classes during incremental updates. Despite simplicity and intuition, we find that such methods suffer from i…

2025

RoomEditor: High-Fidelity Furniture Synthesis with Parameter-Sharing U-Net

NeurIPS 2025poster

Virtual furniture synthesis, a critical task in image composition, aims to seamlessly integrate reference objects into indoor scenes while preserving geometric coherence and visual realism. Despite its significant potential in home design applications, this field remains underexplored due to two maj…

Cited by 0SourcecodeScholar
2025

Self-Convolutional Attention-Based Uncertainty-Aware Network for Single-Image Super-Resolution

ICASSP 2025accepted

Current super-resolution (SR) algorithms rely heavily on annotated data and often ignore the uncertainty in image degradation and features, limiting their real-world application. We propose an uncertainty-aware SR network using a self-convolutional attention mechanism. Our approach focuses on an SR…

Cited by 0SourceScholar
2025

Socialized Coevolution: Advancing a Better World through Cross-Task Collaboration

ICML 2025poster

Traditional machine societies rely on data-driven learning, overlooking interactions and limiting knowledge acquisition from model interplay. To address these issues, we revisit the development of machine societies by drawing inspiration from the evolutionary processes of human societies. Motivated…

2025

Spurious Feature Eraser: Stabilizing Test-Time Adaptation for Vision-Language Foundation Model

AAAI 2025technical

Vision-language foundation models have exhibited remarkable success across a multitude of downstream tasks due to their scalability on extensive image-text paired data. However, these models also display significant limitations when applied to downstream tasks, such as fine-grained image classificat…

2025

TAMT: Temporal-Aware Model Tuning for Cross-Domain Few-Shot Action Recognition

CVPR 2025poster

Going beyond few-shot action recognition (FSAR), cross-domain FSAR (CDFSAR) has attracted recent research interests by solving the domain gap lying in source-to-target transfer learning. Existing CDFSAR methods mainly focus on joint training of source and target data to mitigate the side effect of d…

2025

Task-Gated Multi-Expert Collaboration Network for Degraded Multi-Modal Image Fusion

ICML 2025poster

Multi-modal image fusion aims to integrate complementary information from different modalities to enhance perceptual capabilities in applications such as rescue and security. However, real-world imaging often suffers from degradation issues, such as noise, blur, and haze in visible imaging, as well…

2025

Unknown Text Learning for CLIP-based Few-Shot Open-set Recognition

ICCV 2025poster

Recently, vision-language models (e.g., CLIP) with prompt learning have shown great potential in few-shot learning. However, an open issue remains for the effective extension of CLIP-based models to few-shot open-set recognition (FSOR), which requires classifying known classes and detecting unknown…

2024

AMU-Tuning: Effective Logit Bias for CLIP-based Few-shot Learning

CVPR 2024poster

Recently pre-trained vision-language models (e.g. CLIP) have shown great potential in few-shot learning and attracted a lot of research interest. Although efforts have been made to improve few-shot ability of CLIP key factors on the effectiveness of existing methods have not been well studied limiti…

2024

Ahpatron: A New Budgeted Online Kernel Learning Machine with Tighter Mistake Bound

AAAI 2024technical

In this paper, we study the mistake bound of online kernel learning on a budget. We propose a new budgeted online kernel learning model, called Ahpatron, which significantly improves the mistake bound of previous work and resolves an open problem related to upper bounds of hypothesis space constrain…

2024

Dynamic Brightness Adaptation for Robust Multi-modal Image Fusion

IJCAI 2024poster

Infrared and visible image fusion aim to integrate modality strengths for visually enhanced, informative images. Visible imaging in real-world scenarios is susceptible to dynamic environmental brightness fluctuations, leading to texture degradation. Existing fusion methods lack robustness against su…

2024

Dynamic Sub-graph Distillation for Robust Semi-supervised Continual Learning

AAAI 2024technical

Continual learning (CL) has shown promising results and comparable performance to learning at once in a fully supervised manner. However, CL strategies typically require a large number of labeled samples, making their real-life deployment challenging. In this work, we focus on semi-supervised contin…

2024

Every Node Is Different: Dynamically Fusing Self-Supervised Tasks for Attributed Graph Clustering

AAAI 2024technical

Attributed graph clustering is an unsupervised task that partitions nodes into different groups. Self-supervised learning (SSL) shows great potential in handling this task, and some recent studies simultaneously learn multiple SSL tasks to further boost performance. Currently, different SSL tasks ar…

2024

Exploring Diverse Representations for Open Set Recognition

AAAI 2024technical

Open set recognition (OSR) requires the model to classify samples that belong to closed sets while rejecting unknown samples during test. Currently, generative models often perform better than discriminative models in OSR, but recent studies show that generative models may be computationally infeasi…

2024

ID-like Prompt Learning for Few-Shot Out-of-Distribution Detection

CVPR 2024poster

Out-of-distribution (OOD) detection methods often exploit auxiliary outliers to train model identifying OOD samples especially discovering challenging outliers from auxiliary outliers dataset to improve OOD detection. However they may still face limitations in effectively distinguishing between the…

2024

Out-Of-Distribution Detection with Diversification (Provably)

NeurIPS 2024poster

Out-of-distribution (OOD) detection is crucial for ensuring reliable deployment of machine learning models. Recent advancements focus on utilizing easily accessible auxiliary outliers (e.g., data from the web or other datasets) in training. However, we experimentally reveal that these methods still…

2024

Persistence Homology Distillation for Semi-supervised Continual Learning

NeurIPS 2024poster

Semi-supervised continual learning (SSCL) has attracted significant attention for addressing catastrophic forgetting in semi-supervised data. Knowledge distillation, which leverages data representation and pair-wise similarity, has shown significant potential in preserving information in SSCL. Howev…

2024

Socialized Learning: Making Each Other Better Through Multi-Agent Collaboration

ICML 2024poster

Learning new knowledge frequently occurs in our dynamically changing world, e.g., humans culturally evolve by continuously acquiring new abilities to sustain their survival, leveraging collective intelligence rather than a large number of individual attempts. The effective learning paradigm during c…

2024

Task-Customized Mixture of Adapters for General Image Fusion

CVPR 2024poster

General image fusion aims at integrating important information from multi-source images. However due to the significant cross-task gap the respective fusion mechanism varies considerably in practice resulting in limited performance across subtasks. To handle this problem we propose a novel task-cust…

2024

The Best of Both Worlds: On the Dilemma of Out-of-distribution Detection

NeurIPS 2024poster

Out-of-distribution (OOD) detection is essential for model trustworthiness which aims to sensitively identity semantic OOD samples and robustly generalize for covariate-shifted OOD samples. However, we discover that the superior OOD detection performance of state-of-the-art methods is achieved by se…

2024

What Matters in Graph Class Incremental Learning? An Information Preservation Perspective

NeurIPS 2024poster

Graph class incremental learning (GCIL) requires the model to classify emerging nodes of new classes while remembering old classes. Existing methods are designed to preserve effective information of old models or graph data to alleviate forgetting, but there is no clear theoretical understanding of…

2023

AREA: Adaptive Reweighting via Effective Area for Long-Tailed Classification

ICCV 2023poster

Large-scale data from real-world usually follow a long-tailed distribution (i.e., a few majority classes occupy plentiful training data, while most minority classes have few samples), making the hyperplanes heavily skewed to the minority classes. Traditionally, reweighting is adopted to make the hyp…

Cited by 47PDFcodeScholar
2023

Calibrating Multimodal Learning

ICML 2023oral

Multimodal machine learning has achieved remarkable progress in a wide range of scenarios. However, the reliability of multimodal learning remains largely unexplored. In this paper, through extensive empirical studies, we identify current multimodal classification methods suffer from unreliable pred…

Cited by 20SourcePDFScholar
2023

Exploring and Exploiting Uncertainty for Incomplete Multi-View Classification

CVPR 2023poster

Classifying incomplete multi-view data is inevitable since arbitrary view missing widely exists in real-world applications. Although great progress has been achieved, existing incomplete multi-view methods are still difficult to obtain a trustworthy prediction due to the relatively high uncertainty…

Cited by 30SourcePDFScholar
2023

Fairness-guided Few-shot Prompting for Large Language Models

NeurIPS 2023poster

Large language models have demonstrated surprising ability to perform in-context learning, i.e., these models can be directly applied to solve numerous downstream tasks by conditioning on a prompt constructed by a few input-output examples. However, prior research has shown that in-context learning…

Cited by 82SourcePDFScholar
2023

Multi-Modal Gated Mixture of Local-to-Global Experts for Dynamic Image Fusion

ICCV 2023poster

Infrared and visible image fusion aims to integrate comprehensive information from multiple sources to achieve superior performances on various practical tasks, such as detection, over that of a single modality. However, most existing methods directly combined the texture details and object contrast…

Cited by 47PDFcodeScholar
2023

Provable Dynamic Fusion for Low-Quality Multimodal Data

ICML 2023poster

The inherent challenge of multimodal fusion is to precisely capture the cross-modal correlation and flexibly conduct cross-modal interaction. To fully release the value of each modality and mitigate the influence of low-quality multimodal data, dynamic multimodal fusion emerges as a promising learni…

2023

Public Opinion Field Effect Fusion in Representation Learning for Trending Topics Diffusion

NeurIPS 2023poster

Trending topic diffusion and prediction analysis is an important problem and has been well studied in social networks. Representation learning is an effective way to extract node embeddings, which can help for topic propagation analysis by completing downstream tasks such as link prediction and node…

2023

Reliable and Interpretable Personalized Federated Learning

CVPR 2023poster

Federated learning can coordinate multiple users to participate in data training while ensuring data privacy. The collaboration of multiple agents allows for a natural connection between federated learning and collective intelligence. When there are large differences in data distribution among clien…

Cited by 27SourcePDFScholar
2023

Tuning Pre-trained Model via Moment Probing

ICCV 2023poster

Recently, efficient fine-tuning of large-scale pre-trained models has attracted increasing research interests, where linear probing (LP) as a fundamental module is involved in exploiting the final representations for task-dependent classification. However, most of the existing methods focus on how t…

Cited by 8PDFcodeScholar
2022

DropCov: A Simple yet Effective Method for Improving Deep Architectures

NeurIPS 2022accept

Previous works show global covariance pooling (GCP) has great potential to improve deep architectures especially on visual recognition tasks, where post-normalization of GCP plays a very important role in final performance. Although several post-normalization strategies have been studied, these meth…

2021

Detection, Tracking, and Counting Meets Drones in Crowds: A Benchmark

CVPR 2021poster

To promote the developments of object detection, tracking and counting algorithms in drone-captured videos, we construct a benchmark with a new drone-captured large-scale dataset, named as DroneCrowd, formed by 112 video clips with 33,600 HD frames in various scenarios. Notably, we annotate 20,800 p…

Cited by 133PDFcodeScholar
2021

Multi-View Information-Bottleneck Representation Learning

AAAI 2021technical

In real-world applications, clustering or classification can usually be improved by fusing information from different views. Therefore, unsupervised representation learning on multi-view data becomes a compelling topic in machine learning. In this paper, we propose a novel and flexible unsupervised…

2021

T-SVDNet: Exploring High-Order Prototypical Correlations for Multi-Source Domain Adaptation

ICCV 2021poster

Most existing domain adaptation methods focus on adaptation from only one source domain, however, in practice there are a number of relevant sources that could be leveraged to help improve performance on target domain. We propose a novel approach named T-SVDNet to address the task of Multi-source Do…

Cited by 57PDFcodeScholar
2021

Temporal-attentive Covariance Pooling Networks for Video Recognition

NeurIPS 2021poster

For video recognition task, a global representation summarizing the whole contents of the video snippets plays an important role for the final performance. However, existing video architectures usually generate it by using a simple, global average pooling (GAP) method, which has limited ability to c…

2021

Trustworthy Multimodal Regression with Mixture of Normal-inverse Gamma Distributions

NeurIPS 2021poster

Multimodal regression is a fundamental task, which integrates the information from different sources to improve the performance of follow-up applications. However, existing methods mainly focus on improving the performance and often ignore the confidence of prediction for diverse situations. In this…

2020

ECA-Net: Efficient Channel Attention for Deep Convolutional Neural Networks

CVPR 2020poster

Recently, channel attention mechanism has demonstrated to offer great potential in improving the performance of deep convolutional neural networks (CNNs). However, most existing methods dedicate to developing more sophisticated attention modules for achieving better performance, which inevitably inc…

Cited by 7872PDFcodeScholar
2020

SPL-MLL: Selecting Predictable Landmarks for Multi-Label Learning

ECCV 2020poster

Although significant progress achieved, multi-label classification is still challenging due to the complexity of correlations among different labels. Furthermore, modeling the relationships between input and some (dull) classes further increases the difficulty of accurately predicting all possible l…

Cited by 12SourcePDFScholar
2020

What Deep CNNs Benefit From Global Covariance Pooling: An Optimization Perspective

CVPR 2020poster

Recent works have demonstrated that global covariance pooling (GCP) has the ability to improve performance of deep convolutional neural networks (CNNs) on visual classification task. Despite considerable advance, the reasons on effectiveness of GCP on deep CNNs have not been well studied. In this pa…

Cited by 30PDFcodeScholar
2019

CPM-Nets: Cross Partial Multi-View Networks

NeurIPS 2019spotlight

Despite multi-view learning progressed fast in past decades, it is still challenging due to the difficulty in modeling complex correlation among different views, especially under the context of view missing. To address the challenge, we propose a novel framework termed Cross Partial Multi-View Netwo…

2019

Progressive Image Deraining Networks: A Better and Simpler Baseline

CVPR 2019poster

Along with the deraining performance improvement of deep networks, their structures and learning become more and more complicated and diverse, making it difficult to analyze the contribution of various network modules when developing new deraining networks. To handle this issue, this paper provides…

Cited by 1077PDFcodeScholar
2019

Reciprocal Multi-Layer Subspace Learning for Multi-View Clustering

ICCV 2019poster

Multi-view clustering is a long-standing important research topic, however, remains challenging when handling high-dimensional data and simultaneously exploring the consistency and complementarity of different views. In this work, we present a novel Reciprocal Multi-layer Subspace Learning (RMSL) al…

Cited by 158PDFScholar