← Search

Xin Fan

41 accepted papers

2026

Bridging Human Evaluation to Infrared and Visible Image Fusion

CVPR 2026

Infrared and visible image fusion (IVIF) integrates complementary modalities to enhance scene perception. Current methods predominantly focus on optimizing handcrafted losses and objective metrics, often resulting in fusion outcomes that do not align with human visual preferences. This challenge is

Cited by 0SourcecodeScholar
2026

Dual-Level Hypergraph Generation for Addressing Feature Scarcity in Whole-Slide Image Classification

CVPR 2026

Lymph node metastasis diagnosis in pathological images is a highly challenging four-class classification task, comprising macrometastasis, micrometastasis, isolated tumor cells (ITC), and negative lesions.Unlike conventional classification settings, this four-class scenario simultaneously suffers fr

Cited by 0SourcecodeScholar
2026

RSOD: Reliability-Guided Sonar Image Object Detection with Extremely Limited Labels

AAAI 2026technical

Object detection in sonar images is a key technology in underwater detection systems. Compared to natural images, sonar images contain fewer texture details and are more susceptible to noise, making it difficult for non-experts to distinguish subtle differences between classes. This leads to their i

Cited by 0SourcePDFScholar
2026

Snap2Review: Vision-Grounded Retrieval and Pairwise Preference Alignment for Personalized Reviews Generation

IJCAI 2026

Personalized review generation is vital for e-commerce engagement, yet integrating user-provided images while maintaining distinct user personas remains a critical challenge. Existing methods often struggle to bridge the semantic gap between objective visual signals and subjective linguistic pattern

Cited by 0Scholar
2026

Streaming Diffusion Model for Fast Infrared and Visible Video Fusion

CVPR 2026

Infrared and visible video fusion is pivotal for robust perceptual systems, aiming to synthesize a comprehensive video stream that leverages both thermal resilience and textured details. However, prevailing methods, by treating videos as sequences of independent frames, inherently introduce temporal

Cited by 0SourcecodeScholar
2025

DCEvo: Discriminative Cross-Dimensional Evolutionary Learning for Infrared and Visible Image Fusion

CVPR 2025poster

Infrared and visible image fusion integrates information from distinct spectral bands to enhance image quality by leveraging the strengths and mitigating the limitations of each modality. Existing approaches typically treat image fusion and subsequent high-level tasks as separate processes, resultin…

2025

Enhancing Infrared Vision: Progressive Prompt Fusion Network and Benchmark

NeurIPS 2025poster

We engage in the relatively underexplored task named thermal infrared image enhancement. Existing infrared image enhancement methods primarily focus on tackling individual degradations, such as noise, contrast, and blurring, making it difficult to handle coupled degradations. Meanwhile, all-in-one e…

Cited by 0SourceScholar
2025

Every SAM Drop Counts: Embracing Semantic Priors for Multi-Modality Image Fusion and Beyond

CVPR 2025poster

Multi-modality image fusion, particularly infrared and visible, plays a crucial role in integrating diverse modalities to enhance scene understanding. Although early research prioritized visual quality, preserving fine details and adapting to downstream tasks remains challenging. Recent approaches a…

2025

TextMEF: Text-guided Prompt Learning for Multi-exposure Image Fusion

IJCAI 2025

Multi-exposure image fusion~(MEF) aims to integrate a set of low dynamic range images, producing a single image with a higher dynamic range than either one. Despite significant advancements, current MEF approaches still struggle to handle extremely over- or under-exposed conditions, resulting in uns

Cited by 0SourcePDFScholar
2024

Advancing Generalized Transfer Attack with Initialization Derived Bilevel Optimization and Dynamic Sequence Truncation

IJCAI 2024poster

Transfer attacks generate significant interest for real-world black-box applications by crafting transferable adversarial examples through surrogate models. Whereas, existing works essentially directly optimize the single-level objective w.r.t. the surrogate model, which always leads to poor interpr…

2024

CSCNet: Class-Specified Cascaded Network for Compositional Zero-Shot Learning

ICASSP 2024accepted

Attribute and object (A-O) disentanglement is a fundamental and critical problem for Compositional Zero-shot Learning (CZSL), whose aim is to recognize novel A-O compositions based on foregone knowledge. Existing methods based on disentangled representation learning lose sight of the contextual depe…

Cited by 0SourceScholar
2024

Contourlet Residual for Prompt Learning Enhanced Infrared Image Super-Resolution

ECCV 2024poster

"Image super-resolution (SR) is a critical technique for enhancing image quality, playing a vital role in image enhancement. While recent advancements, notably transformer-based methods, have advanced the field, infrared image SR remains a formidable challenge. Due to the inherent characteristics of…

2024

Depth-Guided Dominant Plane Perception for Unsupervised Homography Estimation

ICASSP 2024accepted

Homography describes the mapping relations of the same plane across views. In scenarios with multiple planes, single homography estimation aims to obtain the optimal solution generated by the largest consistent plane to obey the coplanar constraints. However, existing methods typically consider all…

Cited by 0SourceScholar
2024

Hybrid-Supervised Dual-Search: Leveraging Automatic Learning for Loss-Free Multi-Exposure Image Fusion

AAAI 2024technical

Multi-exposure image fusion (MEF) has emerged as a prominent solution to address the limitations of digital imaging in representing varied exposure levels. Despite its advancements, the field grapples with challenges, notably the reliance on manual designs for network structures and loss functions,…

2024

Towards Robust Image Stitching: An Adaptive Resistance Learning against Compatible Attacks

AAAI 2024technical

Image stitching seamlessly integrates images captured from varying perspectives into a single wide field-of-view image. Such integration not only broadens the captured scene but also augments holistic perception in computer vision applications. Given a pair of captured images, subtle perturbations a…

2024

Trash to Treasure: Low-Light Object Detection via Decomposition-and-Aggregation

AAAI 2024technical

Object detection in low-light scenarios has attracted much attention in the past few years. A mainstream and representative scheme introduces enhancers as the pre-processing for regular detectors. However, because of the disparity in task objectives between the enhancer and detector, this paradigm c…

Cited by 12SourcePDFScholar
2024

Where Elegance Meets Precision: Towards a Compact, Automatic, and Flexible Framework for Multi-modality Image Fusion and Applications

IJCAI 2024poster

Multi-modality image fusion aims to integrate images from multiple sensors, producing an image that is visually appealing and offers more comprehensive information than any single one. To ensure high visual quality and facilitate accurate subsequent perception tasks, previous methods have often casc…

2023

Bi-level Dynamic Learning for Jointly Multi-modality Image Fusion and Beyond

IJCAI 2023poster

Recently, multi-modality scene perception tasks, e.g., image fusion and scene understanding, have attracted widespread attention for intelligent vision systems. However, early efforts always consider boosting a single task unilaterally and neglecting others, seldom investigating their underlying co…

2023

Multi-interactive Feature Learning and a Full-time Multi-modality Benchmark for Image Fusion and Segmentation

ICCV 2023oral

Multi-modality image fusion and segmentation play a vital role in autonomous driving and robotic operation. Early efforts focus on boosting the performance for only one task, e.g., fusion or segmentation, making it hard to reach `Best of Both Worlds'. To overcome this issue, in this paper, we propos…

Cited by 172PDFcodeScholar
2023

Pixels, Regions, and Objects: Multiple Enhancement for Salient Object Detection

CVPR 2023poster

Salient object detection (SOD) aims to mimic the human visual system (HVS) and cognition mechanisms to identify and segment salient objects. However, due to the complexity of these mechanisms, current methods are not perfect. Accuracy and robustness need to be further improved, particularly in compl…

2022

Hierarchical Bilevel Learning with Architecture and Loss Search for Hadamard-based Image Restoration

IJCAI 2022poster

In the past few decades, Hadamard-based image restoration problems (e.g., low-light image enhancement) attract wide concerns in multiple areas related to artificial intelligence. However, existing works mostly focus on heuristically defining architecture and loss by the engineering experiences that…

Cited by 3SourcePDFScholar
2022

ReCoNet: Recurrent Correction Network for Fast and Efficient Multi-Modality Image Fusion

ECCV 2022poster

"Recent advances in deep networks have gained great attention in infrared and visible image fusion (IVIF). Nevertheless, most existing methods are incapable of dealing with slight misalignment on source images and suffer from high computational and spatial expenses. This paper tackles these two crit…

2022

Segment, Magnify and Reiterate: Detecting Camouflaged Objects the Hard Way

CVPR 2022poster

It is challenging to accurately detect camouflaged objects from their highly similar surroundings. Existing methods mainly leverage a single-stage detection fashion, while neglecting small objects with low-resolution fine edges requires more operations than the larger ones. To tackle camouflaged obj…

Cited by 207PDFcodeScholar
2022

Semantic-aware Texture-Structure Feature Collaboration for Underwater Image Enhancement

ICRA 2022poster

Underwater image enhancement has become an attractive topic as a significant technology in marine engi-neering and aquatic robotics. However, the limited number of datasets and imperfect hand-crafted ground truth weaken its robustness to unseen scenarios, and hamper the application to high-level vis…

Cited by 34SourcecodeScholar
2022

Target-Aware Dual Adversarial Learning and a Multi-Scenario Multi-Modality Benchmark To Fuse Infrared and Visible for Object Detection

CVPR 2022oral

This study addresses the issue of fusing infrared and visible images that appear differently for object detection. Aiming at generating an image of high visual quality, previous approaches discover commons underlying the two modalities and fuse upon the common space either by iterative optimization…

Cited by 733PDFcodeScholar
2022

Toward Fast, Flexible, and Robust Low-Light Image Enhancement

CVPR 2022oral

Existing low-light image enhancement techniques are mostly not only difficult to deal with both visual quality and computational efficiency but also commonly invalid in unknown complex scenarios. In this paper, we develop a new Self-Calibrated Illumination (SCI) learning framework for fast, flexible…

Cited by 810PDFcodeScholar
2022

Unsupervised Misaligned Infrared and Visible Image Fusion via Cross-Modality Image Generation and Registration

IJCAI 2022poster

Recent learning-based image fusion methods have marked numerous progress in pre-registered multi-modality data, but suffered serious ghosts dealing with misaligned multi-modality data, due to the spatial deformation and the difficulty narrowing cross-modality discrepancy. To overcome the obstacles,…

2021

GTA-Net: Gradual Temporal Aggregation Network for Fast Video Deraining

ICASSP 2021accepted

Recently, the development of intelligent technology arouses the requirements of high-quality videos. Rain streak is a frequent and inevitable factor to degrade the video. Many researchers have put their energies into eliminating the adverse effects of rainy video. Unfortunately, how to fully utilize…

Cited by 0SourceScholar
2021

Leveraging Line-Point Consistence To Preserve Structures for Wide Parallax Image Stitching

CVPR 2021poster

Generating high-quality stitched images with natural structures is a challenging task in computer vision. In this paper, we succeed in preserving both local and global geometric structures for wide parallax images, while reducing artifacts and distortions. A projective invariant, Characteristic Numb…

Cited by 129PDFcodeScholar
2021

NASA: A Noise-Adaptive and Structure-Aware Learning Framework for Image Deblurring

ICASSP 2021accepted

Image deblurring is a classical low-level visual processing task, which aims to recover a potentially noise-free sharp image from the blurred image. Existing prior-based and learning-based methods usually need to manually set some vital auxiliary components (e.g., noise level). It brings about extre…

Cited by 0SourceScholar
2021

Retinex-Inspired Unrolling With Cooperative Prior Architecture Search for Low-Light Image Enhancement

CVPR 2021poster

Low-light image enhancement plays very important roles in low-level vision areas. Recent works have built a great deal of deep learning models to address this task. However, these approaches mostly rely on significant architecture engineering and suffer from high computational burden. In this paper,…

Cited by 884PDFcodeScholar
2021

Temporal Rain Decomposition with Spatial Structure Guidance for Video Deraining

ICASSP 2021accepted

Recently, removing rain streaks from videos has drawn wide concerns in vision and multimedia communities. But existing works ignore the depicts of image inherent structure and rain location to cause details loss, and their adopted manners of exploiting temporal information are still insufficient. In…

Cited by 0SourceScholar
2020

Bi-level Probabilistic Feature Learning for Deformable Image Registration

IJCAI 2020poster

We address the challenging issue of deformable registration that robustly and efficiently builds dense correspondences between images. Traditional approaches upon iterative energy optimization typically invoke expensive computational load. Recent learning-based methods are able to efficiently predic…

Cited by 0SourcePDFScholar
2020

Image Restoration Via Data-Dependent Proximal Averaged Optimization

ICASSP 2020accepted

Maximum A Posterior (MAP) acts as one of the most popular modeling scheme in image restoration and is usually reduced to a separable optimization model. Unfortunately, it is challenging to establish exact regularization term and the model with complex priors is hard to optimize. In additionally, it…

Cited by 0SourceScholar
2020

Principle-Inspired Multi-Scale Aggregation Network for Extremely Low-Light Image Enhancement

ICASSP 2020accepted

The under-exposure and low-light environments are common to degrade the image-quality with invisible information. To ameliorate this case, a copious of low-light image enhancement methods are developed. However, these existing works are hard to handle extremely low-light conditions with noises, even…

Cited by 0SourceScholar
2020

Sequential Deep Unrolling With Flow Priors For Robust Video Deraining

ICASSP 2020accepted

Video deraining has attracted wide attention since the urgent demand of high-quality video in recent years. The indistinct details and nonideal deraining effects are the most common defects in existing techniques, whose cause lies in the insufficient usage of single-frame image and temporal informat…

Cited by 0SourceScholar
2019

Accurate Monocular 3D Object Detection via Color-Embedded 3D Reconstruction for Autonomous Driving

ICCV 2019poster

In this paper, we propose a monocular 3D object detection framework in the domain of autonomous driving. Unlike previous image-based methods which focus on RGB feature extracted from 2D images, our method solves this problem in the reconstructed 3D space in order to exploit 3D contexts explicitly. T…

Cited by 396PDFScholar
2018

A Bridging Framework for Model Optimization and Deep Propagation

NeurIPS 2018poster

Optimizing task-related mathematical model is one of the most fundamental methodologies in statistic and learning areas. However, generally designed schematic iterations may hard to investigate complex data distributions in real-world applications. Recently, training deep propagations (i.e., network…

Cited by 20SourcePDFScholar
2018

Deep Layer Prior Optimization for Single Image Rain Streaks Removal

ICASSP 2018accepted

Visible distortions caused by rain streaks have significant negative effects on the performance of many vision and learning algorithms. Most of the existing deraining approaches propose to build complex prior models to formulate the appearance of rain streaks. Unfortunately, these human-designed pri…

Cited by 0SourceScholar
2018

Robust Haze Removal Via Joint Deep Transmission and Scene Propagation

ICASSP 2018accepted

Haze is one of the most important factors which reduce the outdoor image quality. Existing approaches often aim to design their models based on principles of hazes. However, even with exactly modeled haze distribution, it is still a challenging task due to factors in real scenario, such as noises, h…

Cited by 0SourceScholar