← Search

Xingyuan Li

17 accepted papers

2026

Bridging Human Evaluation to Infrared and Visible Image Fusion

CVPR 2026

Infrared and visible image fusion (IVIF) integrates complementary modalities to enhance scene perception. Current methods predominantly focus on optimizing handcrafted losses and objective metrics, often resulting in fusion outcomes that do not align with human visual preferences. This challenge is

Cited by 0SourcecodeScholar
2026

DRFusion: Drift-Resilient Temporally Consistent Infrared–Visible Video Fusion

ICML 2026poster

Infrared and visible video fusion is essential for achieving comprehensive perception in dynamic scenes. However, maintaining temporal consistency remains a formidable challenge. Conventional methods relying on optical flow often suffer from geometric rigidity and ghosting artifacts. Moreover, stand…

Cited by 0SourceScholar
2026

Domain Adaptation Guided Infrared and Visible Image Fusion

AAAI 2026technical

Infrared and Visible Image Fusion (IVIF) integrates complementary information from distinct modalities to enhance image quality. However, the effectiveness declines under unseen conditions such as novel weather or scenes, due to domain shifts primarily from variations of data distribution in the vis

Cited by 2SourcePDFScholar
2026

HATIR: Heat-Aware Diffusion for Turbulent Infrared Video Super-Resolution

AAAI 2026technical

Infrared video has been of great interest in visual tasks under challenging environments, but often suffers from severe atmospheric turbulence and compression degradation. Existing video super-resolution (VSR) methods either neglect the inherent modality gap between infrared and visible images or fa

Cited by 0SourcePDFScholar
2026

Toward Real-world Infrared Image Super-Resolution: A Unified Autoregressive Framework and Benchmark Dataset

CVPR 2026

Infrared image super-resolution (IISR) under real-world conditions is a practically significant yet rarely addressed task. Pioneering works are often trained and evaluated on simulated datasets or neglect the intrinsic differences between infrared and visible imaging. In practice, however, real infr

Cited by 0SourcecodeScholar
2026

Uncertainty-Aware Spatial-Frequency Registration and Fusion for Infrared and Visible Images

IJCAI 2026

Infrared and Visible Image Fusion (IVIF) has shown promise in visual tasks under challenging environments, but fusion under unregistered conditions faces inherent misalignments. Current studies to solve them either predict the deformation parameters coarse-to-fine (i.e., coarse registration and fine

Cited by 0Scholar
2026

UniFusion: A Unified Image Fusion Framework with Robust Representation and Source-Aware Preservation

CVPR 2026

Image fusion aims to integrate complementary information from multiple source images to produce a more informative and visually consistent representation, benefiting both human perception and downstream vision tasks. Despite recent progress, most existing fusion methods are designed for specific tas

Cited by 0SourcecodeScholar
2025

DCEvo: Discriminative Cross-Dimensional Evolutionary Learning for Infrared and Visible Image Fusion

CVPR 2025poster

Infrared and visible image fusion integrates information from distinct spectral bands to enhance image quality by leveraging the strengths and mitigating the limitations of each modality. Existing approaches typically treat image fusion and subsequent high-level tasks as separate processes, resultin…

2025

DifIISR: A Diffusion Model with Gradient Guidance for Infrared Image Super-Resolution

CVPR 2025poster

Infrared imaging is essential for autonomous driving and robotic operations as a supportive modality due to its reliable performance in challenging environments. Despite its popularity, the limitations of infrared cameras, such as low spatial resolution and complex degradations, consistently challen…

2025

Efficient Rectified Flow for Image Fusion

NeurIPS 2025poster

Image fusion is a fundamental and important task in computer vision, aiming to combine complementary information from different modalities to fuse images. In recent years, diffusion models have made significant developments in the field of image fusion. However, diffusion models often require comple…

Cited by 0SourceScholar
2025

Toward Automatic Discovery of a Canine Phonetic Alphabet

ACL 2025long

Dogs communicate intelligently but little is known about the phonetic properties of their vocalization communication. For the first time, this paper presents an iterative algorithm inspired by human phonetic discovery, which is based on minimal pairs that determine phonemes by distinguishing differe…

Cited by 0SourcePDFScholar
2024

AEAM3D: Adverse Environment-Adaptive Monocular 3D Object Detection via Feature Extraction Regularization

ICASSP 2024accepted

3D object detection plays a crucial role in intelligent vision systems. Detection in the open world inevitably encounters various adverse scenes while most of existing methods fail in these scenes. To address this issue, this paper proposes a monocular 3D detection model, termed AEAM3D, which effect…

Cited by 0SourceScholar
2024

Adaptive Multi-Exposure Fusion for Enhanced Neural Radiance Fields

ICASSP 2024accepted

Neural Radiance Fields (NeRF) have revolutionized 3D scene modeling and rendering. However, their performance dips when handling images with diverse exposure levels, mainly due to the intricate luminance dynamics. Addressing this, we present an innovative method that proficiently models and renders…

Cited by 0SourceScholar
2024

Contourlet Residual for Prompt Learning Enhanced Infrared Image Super-Resolution

ECCV 2024poster

"Image super-resolution (SR) is a critical technique for enhancing image quality, playing a vital role in image enhancement. While recent advancements, notably transformer-based methods, have advanced the field, infrared image SR remains a formidable challenge. Due to the inherent characteristics of…

2024

Enhancing Neural Radiance Fields with Adaptive Multi-Exposure Fusion: A Bilevel Optimization Approach for Novel View Synthesis

AAAI 2024technical

Neural Radiance Fields (NeRF) have made significant strides in the modeling and rendering of 3D scenes. However, due to the complexity of luminance information, existing NeRF methods often struggle to produce satisfactory renderings when dealing with high and low exposure images. To address this iss…

2024

Phonetic and Lexical Discovery of Canine Vocalization

EMNLP 2024finding

This paper attempts to discover communication patterns automatically within dog vocalizations in a data-driven approach, which breaks the barrier previous approaches that rely on human prior knowledge on limited data. We present a self-supervised approach with HuBERT, enabling the accurate classific…

Cited by 8SourcePDFScholar
2024

Towards Robust Image Stitching: An Adaptive Resistance Learning against Compatible Attacks

AAAI 2024technical

Image stitching seamlessly integrates images captured from varying perspectives into a single wide field-of-view image. Such integration not only broadens the captured scene but also augments holistic perception in computer vision applications. Given a pair of captured images, subtle perturbations a…