← Search

Jinyuan Liu

46 accepted papers

2026

A Hybrid Space Model for Misaligned Multi-modality Image Fusion

AAAI 2026technical

Infrared and visible image fusion aims to integrate complementary information, such as thermal saliency from infrared imagery and fine-grained texture details from visible imagery. However, real-world multi-modal misalignment and geometric deformation often introduce severe artifacts. Most existing

Cited by 1SourcePDFScholar
2026

BiPA: Bilevel Prompt Adaptation for Underwater Instance Segmentation

CVPR 2026

Underwater instance segmentation is essential for fine-grained scene understanding. However, underwater imagery exhibits a strong domain gap from in-air vision due to severe degradation (e.g., turbidity). Consequently, despite its general segmentation ability, SAM degrades sharply underwater. In thi

Cited by 0SourcecodeScholar
2026

Bridging Human Evaluation to Infrared and Visible Image Fusion

CVPR 2026

Infrared and visible image fusion (IVIF) integrates complementary modalities to enhance scene perception. Current methods predominantly focus on optimizing handcrafted losses and objective metrics, often resulting in fusion outcomes that do not align with human visual preferences. This challenge is

Cited by 0SourcecodeScholar
2026

DRFusion: Drift-Resilient Temporally Consistent Infrared–Visible Video Fusion

ICML 2026poster

Infrared and visible video fusion is essential for achieving comprehensive perception in dynamic scenes. However, maintaining temporal consistency remains a formidable challenge. Conventional methods relying on optical flow often suffer from geometric rigidity and ghosting artifacts. Moreover, stand…

Cited by 0SourceScholar
2026

Domain Adaptation Guided Infrared and Visible Image Fusion

AAAI 2026technical

Infrared and Visible Image Fusion (IVIF) integrates complementary information from distinct modalities to enhance image quality. However, the effectiveness declines under unseen conditions such as novel weather or scenes, due to domain shifts primarily from variations of data distribution in the vis

Cited by 0SourcePDFScholar
2026

HATIR: Heat-Aware Diffusion for Turbulent Infrared Video Super-Resolution

AAAI 2026technical

Infrared video has been of great interest in visual tasks under challenging environments, but often suffers from severe atmospheric turbulence and compression degradation. Existing video super-resolution (VSR) methods either neglect the inherent modality gap between infrared and visible images or fa

Cited by 0SourcePDFScholar
2026

HiDRA: Hierarchical Degradation Representation and Adaptation with Generative Priors for Enhancing Infrared Vision

CVPR 2026

Thermal infrared (TIR) imaging enables robust perception in adverse conditions. However, it often suffers from complex degradations (e.g., fixed-pattern noise and low-resolution) due to sensor limitations and environmental dynamics. Existing methods, whether traditional or learning-based, easily fai

Cited by 0SourcecodeScholar
2026

Human-Centric Multi-Exposure Fusion: Benchmark and Bi-level Cognition Distillation Framework

CVPR 2026

Multi-Exposure Fusion (MEF) seeks to generate a single high-quality image from multiple inputs captured at different exposure levels. Despite substantial progress, most existing approaches depend on statistical metrics that poorly reflect human perceptual preferences. Electroencephalography (EEG) pr

Cited by 0SourcecodeScholar
2026

MorphoBall: A Bio-Inspired Transformable Spherical Robot with Dual Terrestrial Gaits and Surface Swimming Capability

ICRA 2026poster

MorphoBall is a bio-inspired, deformable spherical robot designed for multimodal locomotion across terrestrial and aquatic environments. By integrating a dual-mode drive system (spherical rolling and differential-drive) with a morphology-mediated propulsion mechanism, MorphoBall achieves adaptive mo…

Cited by 0Scholar
2026

RSOD: Reliability-Guided Sonar Image Object Detection with Extremely Limited Labels

AAAI 2026technical

Object detection in sonar images is a key technology in underwater detection systems. Compared to natural images, sonar images contain fewer texture details and are more susceptible to noise, making it difficult for non-experts to distinguish subtle differences between classes. This leads to their i

Cited by 0SourcePDFScholar
2026

Streaming Diffusion Model for Fast Infrared and Visible Video Fusion

CVPR 2026

Infrared and visible video fusion is pivotal for robust perceptual systems, aiming to synthesize a comprehensive video stream that leverages both thermal resilience and textured details. However, prevailing methods, by treating videos as sequences of independent frames, inherently introduce temporal

Cited by 0SourcecodeScholar
2026

Taming Generative Diffusion Model for Task-Oriented Infrared Imaging

CVPR 2026

Infrared imaging is essential for perception in harsh environments. However, dynamically coupled degradation factors severely impair visual quality and downstream semantic accuracy. Although generative diffusion models provide strong image restoration priors, high computational cost and physical inc

Cited by 0SourcecodeScholar
2026

Toward Real-world Infrared Image Super-Resolution: A Unified Autoregressive Framework and Benchmark Dataset

CVPR 2026

Infrared image super-resolution (IISR) under real-world conditions is a practically significant yet rarely addressed task. Pioneering works are often trained and evaluated on simulated datasets or neglect the intrinsic differences between infrared and visible imaging. In practice, however, real infr

Cited by 0SourcecodeScholar
2026

Uncertainty-Aware Spatial-Frequency Registration and Fusion for Infrared and Visible Images

IJCAI 2026

Infrared and Visible Image Fusion (IVIF) has shown promise in visual tasks under challenging environments, but fusion under unregistered conditions faces inherent misalignments. Current studies to solve them either predict the deformation parameters coarse-to-fine (i.e., coarse registration and fine

Cited by 0Scholar
2026

UniFusion: A Unified Image Fusion Framework with Robust Representation and Source-Aware Preservation

CVPR 2026

Image fusion aims to integrate complementary information from multiple source images to produce a more informative and visually consistent representation, benefiting both human perception and downstream vision tasks. Despite recent progress, most existing fusion methods are designed for specific tas

Cited by 0SourcecodeScholar
2025

A²RNet: Adversarial Attack Resilient Network for Robust Infrared and Visible Image Fusion

AAAI 2025technical

Infrared and visible image fusion (IVIF) is a crucial technique for enhancing visual performance by integrating unique information from different modalities into one fused image. Exiting methods pay more attention to conducting fusion with undisturbed data, while overlooking the impact of deliberate…

2025

CoA: Towards Real Image Dehazing via Compression-and-Adaptation

CVPR 2025poster

Learning-based image dehazing algorithms have shown remarkable success in synthetic domains. However, real image dehazing is still in suspense due to computational resource constraints and the diversity of real-world scenes. Therefore, there is an urgent need for an algorithm that excels in both eff…

2025

DCEvo: Discriminative Cross-Dimensional Evolutionary Learning for Infrared and Visible Image Fusion

CVPR 2025poster

Infrared and visible image fusion integrates information from distinct spectral bands to enhance image quality by leveraging the strengths and mitigating the limitations of each modality. Existing approaches typically treat image fusion and subsequent high-level tasks as separate processes, resultin…

2025

DEAL: Data-Efficient Adversarial Learning for High-Quality Infrared Imaging

CVPR 2025poster

Thermal imaging is often compromised by dynamic, complex degradations caused by hardware limitations and unpredictable environmental factors. The scarcity of high-quality infrared data, coupled with the challenges of dynamic, intricate degradations, makes it difficult to recover details using exis…

2025

Depth-Supervised Fusion Network for Seamless-Free Image Stitching

NeurIPS 2025poster

Image stitching synthesizes images captured from multiple perspectives into a single image with a broader field of view. The significant variations in object depth often lead to large parallax, resulting in ghosting and misalignment in the stitched results. To address this, we propose a depth-consis…

Cited by 0SourcecodeScholar
2025

DifIISR: A Diffusion Model with Gradient Guidance for Infrared Image Super-Resolution

CVPR 2025poster

Infrared imaging is essential for autonomous driving and robotic operations as a supportive modality due to its reliable performance in challenging environments. Despite its popularity, the limitations of infrared cameras, such as low spatial resolution and complex degradations, consistently challen…

2025

Efficient Rectified Flow for Image Fusion

NeurIPS 2025poster

Image fusion is a fundamental and important task in computer vision, aiming to combine complementary information from different modalities to fuse images. In recent years, diffusion models have made significant developments in the field of image fusion. However, diffusion models often require comple…

Cited by 0SourceScholar
2025

Enhancing Infrared Vision: Progressive Prompt Fusion Network and Benchmark

NeurIPS 2025poster

We engage in the relatively underexplored task named thermal infrared image enhancement. Existing infrared image enhancement methods primarily focus on tackling individual degradations, such as noise, contrast, and blurring, making it difficult to handle coupled degradations. Meanwhile, all-in-one e…

Cited by 0SourceScholar
2025

Every SAM Drop Counts: Embracing Semantic Priors for Multi-Modality Image Fusion and Beyond

CVPR 2025poster

Multi-modality image fusion, particularly infrared and visible, plays a crucial role in integrating diverse modalities to enhance scene understanding. Although early research prioritized visual quality, preserving fine details and adapting to downstream tasks remains challenging. Recent approaches a…

2025

FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

AAAI 2025technical

Large Vision-Language Models (LVLMs) signify a groundbreaking paradigm shift within the Artificial Intelligence (AI) community, extending beyond the capabilities of Large Language Models (LLMs) by assimilating additional modalities (e.g., images). Despite this advancement, the safety of LVLMs remain…

2025

Image Stitching in Adverse Condition: A Bidirectional-Consistency Learning Framework and Benchmark

NeurIPS 2025poster

Deep learning-based image stitching methods have achieved promising performance on conventional stitching datasets. However, real-world scenarios may introduce challenges such as complex weather conditions, illumination variations, and dynamic scene motion, which severely degrade image quality and l…

Cited by 0SourceScholar
2025

Rethinking Reconstruction and Denoising in the Dark: New Perspective, General Architecture and Beyond

CVPR 2025poster

Recently, enhancing image quality in the original RAW domain has garnered significant attention, with denoising and reconstruction emerging as fundamental tasks. Although some works attempt to couple these tasks, they primarily focus on cascade learning while neglecting task associativity within a b…

2025

Task-Specific Information Decomposition for End-to-End Dense Video Captioning

ACL 2025long

Dense video captioning aims to localize events within input videos and generate concise descriptive texts for each event. Advanced end-to-end methods require both tasks to share the same intermediate features that serve as event queries, thereby enabling the mutual promotion of two tasks. However, r…

2025

TextMEF: Text-guided Prompt Learning for Multi-exposure Image Fusion

IJCAI 2025

Multi-exposure image fusion~(MEF) aims to integrate a set of low dynamic range images, producing a single image with a higher dynamic range than either one. Despite significant advancements, current MEF approaches still struggle to handle extremely over- or under-exposed conditions, resulting in uns

Cited by 0SourcePDFScholar
2024

AEAM3D: Adverse Environment-Adaptive Monocular 3D Object Detection via Feature Extraction Regularization

ICASSP 2024accepted

3D object detection plays a crucial role in intelligent vision systems. Detection in the open world inevitably encounters various adverse scenes while most of existing methods fail in these scenes. To address this issue, this paper proposes a monocular 3D detection model, termed AEAM3D, which effect…

Cited by 0SourceScholar
2024

Adaptive Multi-Exposure Fusion for Enhanced Neural Radiance Fields

ICASSP 2024accepted

Neural Radiance Fields (NeRF) have revolutionized 3D scene modeling and rendering. However, their performance dips when handling images with diverse exposure levels, mainly due to the intricate luminance dynamics. Addressing this, we present an innovative method that proficiently models and renders…

Cited by 0SourceScholar
2024

Enhancing Neural Radiance Fields with Adaptive Multi-Exposure Fusion: A Bilevel Optimization Approach for Novel View Synthesis

AAAI 2024technical

Neural Radiance Fields (NeRF) have made significant strides in the modeling and rendering of 3D scenes. However, due to the complexity of luminance information, existing NeRF methods often struggle to produce satisfactory renderings when dealing with high and low exposure images. To address this iss…

2024

Hybrid-Supervised Dual-Search: Leveraging Automatic Learning for Loss-Free Multi-Exposure Image Fusion

AAAI 2024technical

Multi-exposure image fusion (MEF) has emerged as a prominent solution to address the limitations of digital imaging in representing varied exposure levels. Despite its advancements, the field grapples with challenges, notably the reliance on manual designs for network structures and loss functions,…

2024

Improving Cross-Modal Alignment with Synthetic Pairs for Text-Only Image Captioning

AAAI 2024technical

Although image captioning models have made significant advancements in recent years, the majority of them heavily depend on high-quality datasets containing paired images and texts which are costly to acquire. Previous works leverage the CLIP's cross-modal association ability for image captioning, r…

Cited by 11SourcePDFScholar
2024

Local Contrast Prior-Guided Cross Aggregation Model for Effective Infrared Small Target Detection

ICASSP 2024accepted

Infrared small target detection, referring to discovering the precise shapes of dim targets from complex clutter background, has gradually become a hot spot. In recent years, learning-based methods have become the mainstream schemes with high efficiency. However, these methods seldom consider the na…

Cited by 0SourceScholar
2024

Segmentation-Driven Infrared and Visible Image Fusion Via Transformer-Enhanced Architecture Searching

ICASSP 2024accepted

A series of infrared and visible image fusion (IVIF) methods have emerged to improve the performance of segmentation task. However, existing perception-focused IVIF methods take visual effects and semantic information as a unified goal for training, ignoring the task conflicts. Moreover, these metho…

Cited by 0SourceScholar
2024

Towards Robust Image Stitching: An Adaptive Resistance Learning against Compatible Attacks

AAAI 2024technical

Image stitching seamlessly integrates images captured from varying perspectives into a single wide field-of-view image. Such integration not only broadens the captured scene but also augments holistic perception in computer vision applications. Given a pair of captured images, subtle perturbations a…

2024

Trash to Treasure: Low-Light Object Detection via Decomposition-and-Aggregation

AAAI 2024technical

Object detection in low-light scenarios has attracted much attention in the past few years. A mainstream and representative scheme introduces enhancers as the pre-processing for regular detectors. However, because of the disparity in task objectives between the enhancer and detector, this paradigm c…

Cited by 12SourcePDFScholar
2024

Where Elegance Meets Precision: Towards a Compact, Automatic, and Flexible Framework for Multi-modality Image Fusion and Applications

IJCAI 2024poster

Multi-modality image fusion aims to integrate images from multiple sensors, producing an image that is visually appealing and offers more comprehensive information than any single one. To ensure high visual quality and facilitate accurate subsequent perception tasks, previous methods have often casc…

2023

A Homotopy Invariant Based on Convex Dissection Topology and a Distance Optimal Path Planning Algorithm

RA-L 2023

The concept of path homotopy has received widely attention in the field of path planning in recent years. In this letter, a homotopy invariant based on convex dissection for a two-dimensional bounded Euclidean space is developed, which can efficiently encode all homotopy path classes between any two

Cited by 12SourceScholar
2023

Bi-level Dynamic Learning for Jointly Multi-modality Image Fusion and Beyond

IJCAI 2023poster

Recently, multi-modality scene perception tasks, e.g., image fusion and scene understanding, have attracted widespread attention for intelligent vision systems. However, early efforts always consider boosting a single task unilaterally and neglecting others, seldom investigating their underlying co…

2023

CDT-Dijkstra: Fast Planning of Globally Optimal Paths for All Points in 2D Continuous Space

IROS 2023poster

The Dijkstra algorithm is a classic path planning method, which in a discrete graph space, can start from a specified source node and find the shortest path between the source node and all other nodes in the graph. However, to the best of our knowledge, there is no effective method that achieves a f…

Cited by 1SourceScholar
2023

Multi-interactive Feature Learning and a Full-time Multi-modality Benchmark for Image Fusion and Segmentation

ICCV 2023oral

Multi-modality image fusion and segmentation play a vital role in autonomous driving and robotic operation. Early efforts focus on boosting the performance for only one task, e.g., fusion or segmentation, making it hard to reach `Best of Both Worlds'. To overcome this issue, in this paper, we propos…

Cited by 172PDFcodeScholar
2022

ReCoNet: Recurrent Correction Network for Fast and Efficient Multi-Modality Image Fusion

ECCV 2022poster

"Recent advances in deep networks have gained great attention in infrared and visible image fusion (IVIF). Nevertheless, most existing methods are incapable of dealing with slight misalignment on source images and suffer from high computational and spatial expenses. This paper tackles these two crit…

2022

Target-Aware Dual Adversarial Learning and a Multi-Scenario Multi-Modality Benchmark To Fuse Infrared and Visible for Object Detection

CVPR 2022oral

This study addresses the issue of fusing infrared and visible images that appear differently for object detection. Aiming at generating an image of high visual quality, previous approaches discover commons underlying the two modalities and fuse upon the common space either by iterative optimization…

Cited by 733PDFcodeScholar
2022

Unsupervised Misaligned Infrared and Visible Image Fusion via Cross-Modality Image Generation and Registration

IJCAI 2022poster

Recent learning-based image fusion methods have marked numerous progress in pre-registered multi-modality data, but suffered serious ghosts dealing with misaligned multi-modality data, due to the spatial deformation and the difficulty narrowing cross-modality discrepancy. To overcome the obstacles,…