← Search

Pengfei Zhu

43 accepted papers

2026

CtrlFuse: Mask-Prompt Guided Controllable Infrared and Visible Image Fusion

AAAI 2026technical

Infrared and visible image fusion generates all-weather perception-capable images by combining complementary modalities, enhancing environmental awareness for intelligent unmanned systems. Existing methods either focus on pixel-level fusion while overlooking downstream task adaptability or implicitl

Cited by 0SourcePDFScholar
2026

Dream-IF: Dynamic Relative EnhAnceMent for Image Fusion

AAAI 2026technical

Image fusion aims to integrate comprehensive information from images acquired through multiple sources. However, images captured by diverse sensors often encounter various degradations that can negatively affect fusion quality. Traditional fusion methods generally treat image enhancement and fusion

Cited by 0SourcePDFScholar
2026

DroneDINO: Towards Heterogeneous Routed Mixture of Experts for Drone-based Unified Object Detection

ICML 2026oral

Recently, the rapid development of low-altitude aerial applications has driven the need for drone-based unified detectors. In contrast to task-specific detectors that suffer from poor scalability across diverse scenarios, existing unified detectors leverage the Mixture-of-Experts (MoE) architecture …

Cited by 0SourceScholar
2026

MatMart: Material Reconstruction of 3D Objects via Diffusion

CVPR 2026

Applying diffusion models to physically-based material estimation and generation has recently gained prominence. In this paper, we propose MatMart, a novel material reconstruction framework for 3D objects, offering the following advantages. First, MatMart adopts a two-stage reconstruction, starting

Cited by 0SourcecodeScholar
2026

Point Cloud Quantization Through Multimodal Prompting for 3D Understanding

AAAI 2026technical

Vector quantization has emerged as a powerful tool in large-scale multimodal models, unifying heterogeneous representations through discrete token encoding. However, its effectiveness hinges on robust codebook design. Current prototype-based approaches relying on trainable vectors or clustered centr

Cited by 0SourcePDFScholar
2025

Asymmetric Factorized Bilinear Operation for Vision Transformer

ICLR 2025poster

As a core component of Transformer-like deep architectures, a feed-forward network (FFN) for channel mixing is responsible for learning features of each token. Recent works show channel mixing can be enhanced by increasing computational burden or can be slimmed at the sacrifice of performance. Altho…

Cited by 0SourcePDFScholar
2025

Asymmetric Reinforcing Against Multi-Modal Representation Bias

AAAI 2025technical

The strength of multimodal learning lies in its ability to integrate information from various sources, providing rich and comprehensive insights. However, in real-world scenarios, multi-modal systems often face the challenge of dynamic modality contributions, the dominance of different modalities ma…

2025

Decoupled Multi-Predictor Optimization for Inference-Efficient Model Tuning

ICCV 2025poster

Recently, remarkable progress has been made in large-scale pre-trained model tuning, and inference efficiency is becoming more crucial for practical deployment. Early exiting in conjunction with multi-stage predictors, when cooperated with a parameter-efficient fine-tuning strategy, offers a straigh…

2025

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark

ICLR 2025poster

The dynamic imbalance of the fore-background is a major challenge in video object counting, which is usually caused by the sparsity of target objects. This remains understudied in existing works and often leads to severe under-/over-prediction errors. To tackle this issue in video object counting, w…

Cited by 1SourcePDFScholar
2025

Graphs Help Graphs: Multi-Agent Graph Socialized Learning

NeurIPS 2025poster

Graphs in the real world are fragmented and dynamic, lacking collaboration akin to that observed in human societies. Existing paradigms present collaborative information collapse and forgetting, making collaborative relationships poorly autonomous and interactive information insufficient. Moreover,…

Cited by 0SourcecodeScholar
2025

Socialized Coevolution: Advancing a Better World through Cross-Task Collaboration

ICML 2025poster

Traditional machine societies rely on data-driven learning, overlooking interactions and limiting knowledge acquisition from model interplay. To address these issues, we revisit the development of machine societies by drawing inspiration from the evolutionary processes of human societies. Motivated…

2025

Task-Gated Multi-Expert Collaboration Network for Degraded Multi-Modal Image Fusion

ICML 2025poster

Multi-modal image fusion aims to integrate complementary information from different modalities to enhance perceptual capabilities in applications such as rescue and security. However, real-world imaging often suffers from degradation issues, such as noise, blur, and haze in visible imaging, as well…

2024

AMU-Tuning: Effective Logit Bias for CLIP-based Few-shot Learning

CVPR 2024poster

Recently pre-trained vision-language models (e.g. CLIP) have shown great potential in few-shot learning and attracted a lot of research interest. Although efforts have been made to improve few-shot ability of CLIP key factors on the effectiveness of existing methods have not been well studied limiti…

2024

Dynamic Brightness Adaptation for Robust Multi-modal Image Fusion

IJCAI 2024poster

Infrared and visible image fusion aim to integrate modality strengths for visually enhanced, informative images. Visible imaging in real-world scenarios is susceptible to dynamic environmental brightness fluctuations, leading to texture degradation. Existing fusion methods lack robustness against su…

2024

Dynamic Sub-graph Distillation for Robust Semi-supervised Continual Learning

AAAI 2024technical

Continual learning (CL) has shown promising results and comparable performance to learning at once in a fully supervised manner. However, CL strategies typically require a large number of labeled samples, making their real-life deployment challenging. In this work, we focus on semi-supervised contin…

2024

Every Node Is Different: Dynamically Fusing Self-Supervised Tasks for Attributed Graph Clustering

AAAI 2024technical

Attributed graph clustering is an unsupervised task that partitions nodes into different groups. Self-supervised learning (SSL) shows great potential in handling this task, and some recent studies simultaneously learn multiple SSL tasks to further boost performance. Currently, different SSL tasks ar…

2024

Exploring Diverse Representations for Open Set Recognition

AAAI 2024technical

Open set recognition (OSR) requires the model to classify samples that belong to closed sets while rejecting unknown samples during test. Currently, generative models often perform better than discriminative models in OSR, but recent studies show that generative models may be computationally infeasi…

2024

Mitigating Causal Confusion in Vector-Based Behavior Cloning for Safer Autonomous Planning

ICRA 2024poster

The utilization of vector-based deep learning techniques has great prospects in the realm of autonomous driving, particularly in the domains of prediction and planning tasks. However, the application of vector-based backbones for prediction and planning tasks may lead to the occurrence of causal con…

Cited by 1SourceScholar
2024

Persistence Homology Distillation for Semi-supervised Continual Learning

NeurIPS 2024poster

Semi-supervised continual learning (SSCL) has attracted significant attention for addressing catastrophic forgetting in semi-supervised data. Knowledge distillation, which leverages data representation and pair-wise similarity, has shown significant potential in preserving information in SSCL. Howev…

2024

Socialized Learning: Making Each Other Better Through Multi-Agent Collaboration

ICML 2024poster

Learning new knowledge frequently occurs in our dynamically changing world, e.g., humans culturally evolve by continuously acquiring new abilities to sustain their survival, leveraging collective intelligence rather than a large number of individual attempts. The effective learning paradigm during c…

2024

Task-Customized Mixture of Adapters for General Image Fusion

CVPR 2024poster

General image fusion aims at integrating important information from multi-source images. However due to the significant cross-task gap the respective fusion mechanism varies considerably in practice resulting in limited performance across subtasks. To handle this problem we propose a novel task-cust…

2024

What Matters in Graph Class Incremental Learning? An Information Preservation Perspective

NeurIPS 2024poster

Graph class incremental learning (GCIL) requires the model to classify emerging nodes of new classes while remembering old classes. Existing methods are designed to preserve effective information of old models or graph data to alleviate forgetting, but there is no clear theoretical understanding of…

2023

Joint Multi-Level Feature Network for Lightweight Person Re-Identification

ICASSP 2023accepted

Learning fine-grained features is crucial to the performance improvement of person re-identification (Re-ID). Although existing methods have made significant progress, utilizing multi-level information to obtain fine-grained features has not been explored in this field. To alleviate this issue, we p…

Cited by 0SourceScholar
2023

Multi-Modal Gated Mixture of Local-to-Global Experts for Dynamic Image Fusion

ICCV 2023poster

Infrared and visible image fusion aims to integrate comprehensive information from multiple sources to achieve superior performances on various practical tasks, such as detection, over that of a single modality. However, most existing methods directly combined the texture details and object contrast…

Cited by 47PDFcodeScholar
2023

Tuning Pre-trained Model via Moment Probing

ICCV 2023poster

Recently, efficient fine-tuning of large-scale pre-trained models has attracted increasing research interests, where linear probing (LP) as a fundamental module is involved in exploiting the final representations for task-dependent classification. However, most of the existing methods focus on how t…

Cited by 8PDFcodeScholar
2023

Unsupervised Self-Driving Attention Prediction via Uncertainty Mining and Knowledge Embedding

ICCV 2023poster

Predicting attention regions of interest is an important yet challenging task for self-driving systems. Existing methodologies rely on large-scale labeled traffic datasets that are labor-intensive to obtain. Besides, the huge domain gap between natural scenes and traffic scenes in current datasets a…

Cited by 6PDFcodeScholar
2022

Label-Efficient Hybrid-Supervised Learning for Medical Image Segmentation

AAAI 2022technical

Due to the lack of expertise for medical image annotation, the investigation of label-efficient methodology for medical image segmentation becomes a heated topic. Recent progresses focus on the efficient utilization of weak annotations together with few strongly-annotated labels so as to achieve com…

Cited by 32SourcePDFScholar
2021

Detection, Tracking, and Counting Meets Drones in Crowds: A Benchmark

CVPR 2021poster

To promote the developments of object detection, tracking and counting algorithms in drone-captured videos, we construct a benchmark with a new drone-captured large-scale dataset, named as DroneCrowd, formed by 112 video clips with 33,600 HD frames in various scenarios. Notably, we annotate 20,800 p…

Cited by 133PDFcodeScholar
2021

Dynamic Hybrid Relation Exploration Network for Cross-Domain Context-Dependent Semantic Parsing

AAAI 2021technical

Semantic parsing has long been a fundamental problem in natural language processing. Recently, cross-domain context-dependent semantic parsing has become a new focus of research. Central to the problem is the challenge of leveraging contextual information of both natural language queries and databas…

Cited by 61SourcePDFScholar
2021

Multi-View Information-Bottleneck Representation Learning

AAAI 2021technical

In real-world applications, clustering or classification can usually be improved by fusing information from different views. Therefore, unsupervised representation learning on multi-view data becomes a compelling topic in machine learning. In this paper, we propose a novel and flexible unsupervised…

2020

ECA-Net: Efficient Channel Attention for Deep Convolutional Neural Networks

CVPR 2020poster

Recently, channel attention mechanism has demonstrated to offer great potential in improving the performance of deep convolutional neural networks (CNNs). However, most existing methods dedicate to developing more sophisticated attention modules for achieving better performance, which inevitably inc…

Cited by 7872PDFcodeScholar
2020

SPL-MLL: Selecting Predictable Landmarks for Multi-Label Learning

ECCV 2020poster

Although significant progress achieved, multi-label classification is still challenging due to the complexity of correlations among different labels. Furthermore, modeling the relationships between input and some (dull) classes further increases the difficulty of accurately predicting all possible l…

Cited by 12SourcePDFScholar
2020

Spatial Attention Pyramid Network for Unsupervised Domain Adaptation

ECCV 2020poster

Unsupervised domain adaptation is critical in various computer vision tasks, such as object detection, instance segmentation, and semantic segmentation, which aims to alleviate performance degradation caused by domain-shift. Most of previous methods rely on a single-mode distribution of source and t…

Cited by 137SourcePDFScholar
2019

Progressive Image Deraining Networks: A Better and Simpler Baseline

CVPR 2019poster

Along with the deraining performance improvement of deep networks, their structures and learning become more and more complicated and diverse, making it difficult to analyze the contribution of various network modules when developing new deraining networks. To handle this issue, this paper provides…

Cited by 1077PDFcodeScholar