← Search

Chunhong Pan

36 accepted papers

2026

IF-Bench: Benchmarking and Enhancing MLLMs for Infrared Images with Generative Visual Prompting

CVPR 2026

Recent advances in multimodal large language models (MLLMs) have led to impressive progress across various benchmarks. However, their capability in understanding infrared images remains unexplored. To address this gap, we introduce **IF-Bench**, the first high-quality benchmark designed for evaluati

Cited by 0SourcecodeScholar
2025

Diffusion Model as a Noise-Aware Latent Reward Model for Step-Level Preference Optimization

NeurIPS 2025poster

Preference optimization for diffusion models aims to align them with human preferences for images. Previous methods typically use Vision-Language Models (VLMs) as pixel-level reward models to approximate human preferences. However, when used for step-level preference optimization, these models face…

Cited by 0SourcecodeScholar
2025

UNIP: Rethinking Pre-trained Attention Patterns for Infrared Semantic Segmentation

ICLR 2025poster

Pre-training techniques significantly enhance the performance of semantic segmentation tasks with limited training data. However, the efficacy under a large domain gap between pre-training (e.g. RGB) and fine-tuning (e.g. infrared) remains underexplored. In this study, we first benchmark the infrare…

2024

Weak Distribution Detectors Lead to Stronger Generalizability of Vision-Language Prompt Tuning

AAAI 2024technical

We propose a generalized method for boosting the generalization ability of pre-trained vision-language models (VLMs) while fine-tuning on downstream few-shot tasks. The idea is realized by exploiting out-of-distribution (OOD) detection to predict whether a sample belongs to a base distribution or a…

2022

AME: Attention and Memory Enhancement in Hyper-Parameter Optimization

CVPR 2022poster

Training Deep Neural Networks (DNNs) is inherently subject to sensitive hyper-parameters and untimely feedbacks of performance evaluation. To solve these two difficulties, an efficient parallel hyper-parameter optimization model is proposed under the framework of Deep Reinforcement Learning (DRL). T…

Cited by 5PDFScholar
2022

Learning from the Target: Dual Prototype Network for Few Shot Semantic Segmentation

AAAI 2022technical

Due to the scarcity of annotated samples, the diversity between support set and query set becomes the main obstacle for few shot semantic segmentation. Most existing prototype-based approaches only exploit the prototype from the support feature and ignore the information from the query sample, faili…

Cited by 22SourcePDFScholar
2022

Stereo Depth Estimation with Echoes

ECCV 2022poster

"Stereo depth estimation is particularly amenable to local textured regions while echoes have good depth estimations for global textureless regions, thus the two modalities complement each other. Motivated by the reciprocal relationship between both modalities, in this paper, we propose an end-to-en…

2021

Differentiable Convolution Search for Point Cloud Processing

ICCV 2021poster

Exploiting convolutional neural networks for point cloud processing is quite challenging, due to the inherent irregular distribution and discrete shape representation of point clouds. To address these problems, many handcrafted convolution variants have sprung up in recent years. Though with elabora…

Cited by 10PDFScholar
2021

Knowledge Mining and Transferring for Domain Adaptive Object Detection

ICCV 2021poster

With the thriving of deep learning, CNN-based object detectors have made great progress in the past decade. However, the domain gap between training and testing data leads to a prominent performance degradation and thus hinders their application in the real world. To alleviate this problem, Knowledg…

Cited by 69PDFcodeScholar
2021

Ltaf-Net: Learning Task-Aware Adaptive Features and Refining Mask for Few-Shot Semantic Segmentation

ICASSP 2021accepted

Few shot segmentation is a newly-developing and challenging computer vision task which is only provided with few labeled samples of the novel class. Some recent works on this problem focus more on how to design an effective comparison module but ignore how to extract the features passed to compare.…

Cited by 0SourceScholar
2021

Reinforcement Stacked Learning with Semantic-Associated Attention for Visual Question Answering

ICASSP 2021accepted

The task of visual question answering (VQA) is to generate an answer for a question according to the content of an image being asked. In this process, the critical problems of effectively embedding the question feature and image feature as well as transforming the features to the prediction of answe…

Cited by 0SourceScholar
2020

AugFPN: Improving Multi-Scale Feature Learning for Object Detection

CVPR 2020poster

Current state-of-the-art detectors typically exploit feature pyramid to detect objects at different scales. Among them, FPN is one of the representative works that build a feature pyramid by multi-scale features summation. However, the design defects behind prevent the multi-scale features from bein…

Cited by 589PDFcodeScholar
2020

Decoupled Representation Learning for Skeleton-Based Gesture Recognition

CVPR 2020poster

Skeleton-based gesture recognition is very challenging, as the high-level information in gesture is expressed by a sequence of complexly composite motions. Previous works often learn all the motions with a single model. In this paper, we propose to decouple the gesture into hand posture variations a…

Cited by 94PDFScholar
2020

Learning Where to Focus for Efficient Video Object Detection

ECCV 2020poster

Transferring existing image-based detectors to the video is non-trivial since the quality of frames is always deteriorated by part occlusion, rare pose, and motion blur. Previous approaches exploit to propagate and aggregate features across video frames by using optical flow-warping. However, direct…

2020

PackDet: Packed Long-Head Object Detector

ECCV 2020poster

State-of-the-art object detectors exploit multi-branch structure and predict objects at several different scales, although substantially boosted accuracy is acquired, low efficiency is inevitable as fragmented structure is hardware unfriendly. To solve this issue, we propose a packing operator (Pack…

2019

DATA: Differentiable ArchiTecture Approximation

NeurIPS 2019poster

Neural architecture search (NAS) is inherently subject to the gap of architectures during searching and validating. To bridge this gap, we develop Differentiable ArchiTecture Approximation (DATA) with an Ensemble Gumbel-Softmax (EGS) estimator to automatically approximate architectures during search…

2019

DensePoint: Learning Densely Contextual Representation for Efficient Point Cloud Processing

ICCV 2019poster

Point cloud processing is very challenging, as the diverse shapes formed by irregular points are often indistinguishable. A thorough grasp of the elusive shape requires sufficiently contextual semantic information, yet few works devote to this. Here we propose DensePoint, a general architecture to l…

Cited by 368PDFcodeScholar
2019

Progressive Sparse Local Attention for Video Object Detection

ICCV 2019poster

Transferring image-based object detectors to the domain of videos remains a challenging problem. Previous efforts mostly exploit optical flow to propagate features across frames, aiming to achieve a good trade-off between accuracy and efficiency. However, introducing an extra model to estimate optic…

Cited by 113PDFScholar
2019

Relation-Shape Convolutional Neural Network for Point Cloud Analysis

CVPR 2019oral

Point cloud analysis is very challenging, as the shape implied in irregular points is difficult to capture. In this paper, we propose RS-CNN, namely, Relation-Shape Convolutional Neural Network, which extends regular grid CNN to irregular configuration for point cloud analysis. The key to RS-CNN is…

Cited by 1206PDFcodeScholar
2018

Exploiting Vector Fields for Geometric Rectification of Distorted Document Images

ECCV 2018poster

This paper proposes a segment-free method for geometric rectification of a distorted document image captured by a hand-held camera. The method can recover the 3D page shape by exploiting the intrinsic vector fields of the image. Based on the assumption that the curled page shape is a general cylindr…

Cited by 29SourcePDFScholar
2018

Fast Variational Level Set Based Image Segmentation via Two-Scale Filtering Model

ICASSP 2018accepted

One major difficulty in medical image segmentation is intensity inhomogeneity, which manifests itself with a slow intensity variation over the whole image domain. Recently, a local binary fitting (LBF) model has been proposed to solve this problem within level set segmentation framework. However, th…

Cited by 0SourceScholar
2018

Structure-Aware Convolutional Neural Networks

NeurIPS 2018poster

Convolutional neural networks (CNNs) are inherently subject to invariable filters that can only aggregate local inputs with the same topological structures. It causes that CNNs are allowed to manage data with Euclidean or grid-like structures (e.g., images), not ones with non-Euclidean or graph stru…

2017

Learning deep vector regression model for no-reference image quality assessment

ICASSP 2017accepted

The goal of no-reference image quality assessment (NR-IQA) is to estimate human perceived image quality without access to either reference image or prior knowledge about distortion type. Previous approaches for this problem are typically based on a regression framework that maps the image features d…

Cited by 0SourceScholar
2017

RoDLSR: Robust discriminative least squares regression model for multi-category classification

ICASSP 2017accepted

Discriminative least squares regression (DLSR) is a simple yet effective method for multi-class classification. One problem of DLSR is that it is lack of robustness to outliers. In order to tackle this difficulty, in this paper, we propose a novel Robust DLSR (RoDLSR) model. The core idea behind RoD…

Cited by 0SourceScholar
2016

Fine-structured object segmentation via edge-guided graph cut with interaction simplification

ICASSP 2016accepted

Fine-structured object segmentation is a challenging problem in object segmentation community. There are mainly two difficulties that can seriously degrade the segmentation quality: 1) insufficient interactions on fine structures due to the high demand of time and manual efforts, and 2) shrinking bi…

Cited by 0SourceScholar
2015

Extraction of Virtual Baselines From Distorted Document Images Using Curvilinear Projection

ICCV 2015poster

The baselines of a document page are a set of virtual horizontal and parallel lines, to which the printed contents of document, e.g., text lines, tables or inserted photos, are aligned. Accurate baseline extraction is of great importance in the geometric correction of curved document images. In this…

Cited by 17PDFScholar