← Search

Jiayi Ma

77 accepted papers

2026

Diff-NAT: Better Naturalistic and Aggressive Adversarial Attacks via Class-Optimized Diffusion for Object Detection

AAAI 2026technical

Recent advances in naturalistic physical adversarial patch generation show great promise in protecting personal privacy against detector-based malicious surveillance while remaining inconspicuous to human observers. In this work, we present the first systematic categorization and in-depth re-examina

Cited by 0SourcePDFScholar
2026

GeoMoE: Divide-and-Conquer Motion Field Modeling with Mixture-of-Experts for Two-View Geometry

AAAI 2026technical

Recent progress in two-view geometry increasingly emphasizes enforcing smoothness and global consistency priors when estimating motion fields between pairs of images. However, in complex real-world scenes, characterized by extreme viewpoint and scale changes as well as pronounced depth discontinuiti

Cited by 0SourcePDFScholar
2026

MagicFuse: Single Image Fusion for Visual and Semantic Reinforcement

CVPR 2026

This paper focuses on a highly practical scenario: how to continue benefiting from the advantages of multi-modal image fusion under harsh conditions when only visible imaging sensors are available. To achieve this goal, we propose a novel concept of single image fusion, which extends conventional da

Cited by 0SourcecodeScholar
2026

Probabilistic Deformation Consistency for Unsupervised Shape Matching

AAAI 2026technical

In this paper, we propose a novel unsupervised shape matching framework based on probabilistic deformation consistency in the spectral domain, termed as PDCMatch. Axiomatic optimization methods suffer from expensive geodesic distance calculations and vulnerability to local optima, and learning-based

Cited by 0SourcePDFScholar
2026

ReCoFuse: Ultra-Robust Image Fusion via Restorative Multi-Modal Diffusion Reciprocal Coupling

CVPR 2026

Existing methods following the integrated hard-regression or decoupling optimization paradigms exhibit limited fusion performance under complex degradations. To address these paradigm-level shortcomings, we propose ReCoFuse, an ultra-robust image fusion framework based on restorative multi-modal dif

Cited by 0SourcecodeScholar
2026

Residual Diffusion Bridge Model for Image Restoration

CVPR 2026

Diffusion bridge models establish probabilistic paths between arbitrary paired distributions and exhibit great potential for universal image restoration. Most existing methods merely treat them as simple variants of stochastic interpolants, lacking a unified analytical perspective. Besides, they ind

Cited by 0SourcecodeScholar
2026

Robust Fusion Controller: Degradation-Aware Image Fusion with Fine-Grained Language Instructions

AAAI 2026technical

Current image fusion methods struggle to adapt to real-world environments encompassing diverse degradations with spatially varying characteristics. To address this challenge, we propose a robust fusion controller (RFC) capable of achieving degradation-aware image fusion through fine-grained language

Cited by 0SourcePDFScholar
2026

SAG-GNN: Semantic-Aware Guided GNN for Descriptor-Free 2D-3D Matching

CVPR 2026

Image-to-point cloud matching (2D-3D matching) establishes accurate correspondences between image keypoints and 3D points for 6-DoF camera pose estimation. Existing methods either suffer from poor generalization due to scene-specific coordinate regression requiring per-scene retraining, or incur hig

Cited by 0SourcecodeScholar
2026

SGPFeat: Semantic and Geometric Priors for Multi-modal Image Matching

AAAI 2026technical

Multi-modal image matching is a fundamental task in multi-view and multi-modal image processing. Its key challenge lies in extracting features that remain consistent despite drastic appearance variations across modalities. However, the learning of the feature is hindered by the scarcity and the inac

Cited by 0SourcePDFScholar
2026

SpeciFuse: Learning Degradation-Type Specificity for Robust Infrared and Visible Image Fusion Under Composite Degradations

IJCAI 2026

Existing degradation-resistant infrared-visible image fusion methods struggle to effectively handle composite degradations, where multiple degradation types exhibit intricate coupling and mutual interference. To address this challenge, we propose SpeciFuse, an infrared-visible image fusion network t

Cited by 0Scholar
2026

VideoFusion: A Spatio-Temporal Collaborative Network for Multi-modal Video Fusion

CVPR 2026

Compared to images, videos better reflect real-world acquisition and possess valuable temporal cues. However, existing multi-sensor fusion research predominantly integrates complementary context from multiple images rather than videos due to the scarcity of large-scale multi-sensor video datasets, l

Cited by 0SourcecodeScholar
2025

Adapting Dense Matching for Homography Estimation with Grid-based Acceleration

CVPR 2025poster

Current deep homography estimation methods are typically constrained to processing low-resolution image pairs due to network architecture and computational limitations. For high-resolution images, downsampling is often required, which can greatly degrade estimation accuracy. In contrast, image match…

2025

ArgMatch: Adaptive Refinement Gathering for Efficient Dense Matching

ICCV 2025poster

Establishing dense correspondences is crucial yet computationally demanding in multi-view tasks. Although coarse-to-fine schemes mitigate computational costs, their efficiency remains limited by the substantial demands of heavy feature extractors and global matchers. In this paper, we propose Adapti…

2025

Balancing Task-invariant Interaction and Task-specific Adaptation for Unified Image Fusion

ICCV 2025poster

Unified image fusion aims to integrate complementary information from multi-source images, enhancing image quality through a unified framework applicable to diverse fusion tasks. While treating all fusion tasks as a unified problem facilitates task-invariant knowledge sharing, it often overlooks tas…

2025

CoMatch: Dynamic Covisibility-Aware Transformer for Bilateral Subpixel-Level Semi-Dense Image Matching

ICCV 2025poster

This prospective study proposes CoMatch, a novel semi-dense image matcher with dynamic covisibility awareness and bilateral subpixel accuracy. Firstly, observing that modeling context interaction over the entire coarse feature map elicits highly redundant computation due to the neighboring represent…

2025

ControlFusion: A Controllable Image Fusion Network with Language-Vision Degradation Prompts

NeurIPS 2025oral

Current image fusion methods struggle with real-world composite degradations and lack the flexibility to accommodate user-specific needs. To address this, we propose ControlFusion, a controllable fusion network guided by language-vision prompts that adaptively mitigates composite degradations. On th…

Cited by 0SourceScholar
2025

DGSolver: Diffusion Generalist Solver with Universal Posterior Sampling for Image Restoration

NeurIPS 2025poster

Diffusion models have achieved remarkable progress in universal image restoration. However, existing methods perform naive inference in the reverse process, which leads to cumulative errors under limited sampling steps and large step intervals. Moreover, they struggle to balance the commonality of d…

Cited by 0SourcecodeScholar
2025

DeMo: Deep Motion Field Consensus with Learnable Kernels for Two-view Correspondence Learning

AAAI 2025technical

As a long-range prior, motion consensus essentially forces the overall spatial transformation between a pair of images to be smooth and consistent, which is naturally well-suited for two-view correspondence learning. However, such precious property remains under-explored by most existing studies due…

2025

Deep Adaptive Unfolded Network via Spatial Morphology Stripping and Spectral Filtration for Pan-sharpening

ICCV 2025poster

In the field of pan-sharpening, existing deep methods are hindered in deepening cross-modal complementarity in the intermediate feature, and lack effective strategies to harness the network entirety for optimal solutions, exhibiting limited feasibility and interpretability due to their black-box des…

2025

Deno-IF: Unsupervised Noisy Visible and Infrared Image Fusion Method

NeurIPS 2025spotlight

Most image fusion methods are designed for ideal scenarios and struggle to handle noise. Existing noise-aware fusion methods are supervised and heavily rely on constructed paired data, limiting performance and generalization. This paper proposes a novel unsupervised noisy visible and infrared image…

Cited by 0SourcecodeScholar
2025

End-to-End Entity-Predicate Association Reasoning for Dynamic Scene Graph Generation

ICCV 2025poster

Dynamic Scene Graph Generation (DSGG) aims to comprehensively understand videos by abstracting them into visual triplets <subject, predicate, object>. Most existing methods focus on capturing temporal dependencies, but overlook crucial visual relationship dependencies between entities and predicates…

2025

HyperGCT: A Dynamic Hyper-GNN-Learned Geometric Constraint for 3D Registration

ICCV 2025poster

Geometric constraints between feature matches are critical in 3D point cloud registration problems. Existing approaches typically model unordered matches as a consistency graph and sample consistent matches to generate hypotheses. However, explicit graph construction introduces noise, posing great c…

2025

LUT-Fuse: Towards Extremely Fast Infrared and Visible Image Fusion via Distillation to Learnable Look-Up Tables

ICCV 2025poster

Current advanced research on infrared and visible image fusion primarily focuses on improving fusion performance, often neglecting the applicability on real-time fusion devices. In this paper, we propose a novel approach that towards extremely fast fusion via distillation to learnable lookup tables…

2025

Matching While Perceiving: Enhance Image Feature Matching with Applicable Semantic Amalgamation

AAAI 2025technical

Image feature matching is a cardinal problem in computer vision, aiming to establish accurate correspondences between two-view images. Existing methods are constrained by the performance of feature extractors and struggle to capture local information affected by sparse texture or occlusions. Recogni…

2025

Multi-Shape Matching with Cycle Consistency Basis via Functional Maps

AAAI 2025technical

Multi-shape matching is a central problem in various applications of computer vision and graphics, where cycle consistency constraints play a pivotal role. For this issue, we propose a novel and efficient approach that models multi-shapes as directed graphs for two-stage optimization, i.e., optimizi…

2025

Multimodal Image Matching Based on Cross-Modality Completion Pre-training

IJCAI 2025

The differences in imaging devices cause multimodal images to have modal differences and geometric distortions, complicating the matching task. Deep learning-based matching methods struggle with multimodal images due to the lack of large annotated multimodal datasets. To address these challenges, we

Cited by 0SourcePDFScholar
2025

Projection-Manifold Regularized Latent Diffusion for Robust General Image Fusion

NeurIPS 2025poster

This study proposes PDFuse, a robust, general training-free image fusion framework built on pre-trained latent diffusion models with projection–manifold regularization. By redefining fusion as a diffusion inference process constrained by multiple source images, PDFuse can adapt to varied image modal…

Cited by 0SourcecodeScholar
2025

Robust Test-Time Adaptation for Single Image Denoising Using Deep Gaussian Prior

ICCV 2025poster

Gaussian denoising often serves as the initiation of research in the field of image denoising, owing to its prevalence and intriguing properties. However, deep Gaussian denoiser typically generalizes poorly to other types of noises, such as Poisson noise and real-world noise. In this paper, we revea…

2024

A Robust Mutual-Reinforcing Framework for 3D Multi-Modal Medical Image Fusion Based on Visual-Semantic Consistency

AAAI 2024technical

This work proposes a robust 3D medical image fusion framework to establish a mutual-reinforcing mechanism between visual fusion and lesion segmentation, achieving their double improvement. Specifically, we explore the consistency between vision and semantics by sharing feature fusion modules. Throug…

2024

Aux-NAS: Exploiting Auxiliary Labels with Negligibly Extra Inference Cost

ICLR 2024poster

We aim at exploiting additional auxiliary labels from an independent (auxiliary) task to boost the primary task performance which we focus on, while preserving a single task inference cost of the primary task. While most existing auxiliary learning methods are optimization-based relying on loss weig…

2024

CLIFF: Continual Latent Diffusion for Open-Vocabulary Object Detection

ECCV 2024oral

"Open-vocabulary object detection (OVD) utilizes image-level cues to expand the linguistic space of region proposals, thereby facilitating the detection of diverse novel classes. Recent works adapt CLIP embedding by minimizing the object-image and object-text discrepancy combinatorially in a discrim…

2024

Cross-Scale Domain Adaptation with Comprehensive Information for Pansharpening

IJCAI 2024poster

Deep learning-based pansharpening methods typically use simulated data at the reduced-resolution scale for training. It limits their performance when generalizing the trained model to the full-resolution scale due to incomprehensive information utilization of panchromatic (PAN) images at the full-re…

2024

DeMatch: Deep Decomposition of Motion Field for Two-View Correspondence Learning

CVPR 2024poster

Two-view correspondence learning has recently focused on considering the coherence and smoothness of the motion field between an image pair. Dominant schemes include controlling the complexity of the field function with regularization or smoothing the field with local filters but the former suffers…

2024

Deep Unfolded Network with Intrinsic Supervision for Pan-Sharpening

AAAI 2024technical

Existing deep pan-sharpening methods lack the learning of complementary information between PAN and MS modalities in the intermediate layers, and exhibit low interpretability due to their black-box designs. To this end, an interpretable deep unfolded network with intrinsic supervision for pan-sharpe…

2024

Dispel Darkness for Better Fusion: A Controllable Visual Enhancer based on Cross-modal Conditional Adversarial Learning

CVPR 2024poster

We propose a controllable visual enhancer named DDBF which is based on cross-modal conditional adversarial learning and aims to dispel darkness and achieve better visible and infrared modalities fusion. Specifically a guided restoration module (GRM) is firstly designed to enhance weakened informatio…

2024

Locality Preserving Refinement for Shape Matching with Functional Maps

AAAI 2024technical

In this paper, we address the nonrigid shape matching with outliers by a novel and effective pointwise map refinement method, termed Locality Preserving Refinement. For accurate pointwise conversion from a given functional map, our method formulates a two-step procedure. Firstly, starting with noisy…

2024

MRFS: Mutually Reinforcing Image Fusion and Segmentation

CVPR 2024poster

This paper proposes a coupled learning framework to break the performance bottleneck of infrared-visible image fusion and segmentation called MRFS. By leveraging the intrinsic consistency between vision and semantics it emphasizes mutual reinforcement rather than treating these tasks as separate iss…

2024

ResMatch: Residual Attention Learning for Feature Matching

AAAI 2024technical

Attention-based graph neural networks have made great progress in feature matching. However, the literature lacks a comprehensive understanding of how the attention mechanism operates for feature matching. In this paper, we rethink cross- and self-attention from the viewpoint of traditional feature…

2024

SDGMNet: Statistic-Based Dynamic Gradient Modulation for Local Descriptor Learning

AAAI 2024technical

Rescaling the backpropagated gradient of contrastive loss has made significant progress in descriptor learning. However, current gradient modulation strategies have no regard for the varying distribution of global gradients, so they would suffer from changes in training phases or datasets. In this p…

2024

Text-DiFuse: An Interactive Multi-Modal Image Fusion Framework based on Text-modulated Diffusion Model

NeurIPS 2024spotlight

Existing multi-modal image fusion methods fail to address the compound degradations presented in source images, resulting in fusion images plagued by noise, color bias, improper exposure, etc. Additionally, these methods often overlook the specificity of foreground objects, weakening the salience of…

2024

Text-IF: Leveraging Semantic Text Guidance for Degradation-Aware and Interactive Image Fusion

CVPR 2024poster

Image fusion aims to combine information from different source images to create a comprehensively representative image. Existing fusion methods are typically helpless in dealing with degradations in low-quality source images and non-interactive to multiple subjective and objective needs. To solve th…

2024

Unmixing Before Fusion: A Generalized Paradigm for Multi-Source-based Hyperspectral Image Synthesis

CVPR 2024poster

In the realm of AI data serves as a pivotal resource. Real-world hyperspectral images (HSIs) bearing wide spectral characteristics are particularly valuable. However the acquisition of HSIs is always costly and time-intensive resulting in a severe data-thirsty issue in HSI research and applications.…

2023

Diff-Retinex: Rethinking Low-light Image Enhancement with A Generative Diffusion Model

ICCV 2023poster

In this paper, we rethink the low-light image enhancement task and propose a physically explainable and generative diffusion model for low-light image enhancement, termed as Diff-Retinex. We aim to integrate the advantages of the physical model and the generative network. Furthermore, we hope to sup…

Cited by 141PDFScholar
2023

LSTFE-Net:Long Short-Term Feature Enhancement Network for Video Small Object Detection

CVPR 2023poster

Video small object detection is a difficult task due to the lack of object information. Recent methods focus on adding more temporal information to obtain more potent high-level features, which often fail to specify the most vital information for small objects, resulting in insufficient or inappropr…

2023

Robust and Scalable Gaussian Process Regression and Its Applications

CVPR 2023poster

This paper introduces a robust and scalable Gaussian process regression (GPR) model via variational learning. This enables the application of Gaussian processes to a wide range of real data, which are often large-scale and contaminated by outliers. Towards this end, we employ a mixture likelihood mo…

2023

Sparsely Annotated Semantic Segmentation With Adaptive Gaussian Mixtures

CVPR 2023poster

Sparsely annotated semantic segmentation (SASS) aims to learn a segmentation model by images with sparse labels (i.e., points or scribbles). Existing methods mainly focus on introducing low-level affinity or generating pseudo labels to strengthen supervision, while largely ignoring the inherent rela…

2023

U-Match: Two-view Correspondence Learning with Hierarchy-aware Local Context Aggregation

IJCAI 2023poster

Local context capturing has become the core factor for achieving leading performance in two-view correspondence learning. Recent advances have devised various local context extractors whereas typically adopting explicit neighborhood relation modeling that is restricted and inflexible. To address thi…

2023

Unsupervised Multi-Exposure Image Fusion Breaking Exposure Limits via Contrastive Learning

AAAI 2023technical

This paper proposes an unsupervised multi-exposure image fusion (MEF) method via contrastive learning, termed as MEF-CL. It breaks exposure limits and performance bottleneck faced by existing methods. MEF-CL firstly designs similarity constraints to preserve contents in source images. It eliminates…

2022

Coherent Point Drift Revisited for Non-Rigid Shape Matching and Registration

CVPR 2022poster

In this paper, we explore a new type of extrinsic method to directly align two geometric shapes with point-to-point correspondences in ambient space by recovering a deformation, which allows more continuous and smooth maps to be obtained. Specifically, the classic coherent point drift is revisited a…

Cited by 17PDFScholar
2022

Fusion from Decomposition: A Self-Supervised Decomposition Approach for Image Fusion

ECCV 2022poster

"Image fusion is famous as an alternative solution to generate one high-quality image from multiple images in addition to image restoration from a single degraded image. The essence of image fusion is to integrate complementary information or best parts from source images. The current fusion methods…

Cited by 140SourcePDFScholar
2022

Hierarchical Memory Learning for Fine-Grained Scene Graph Generation

ECCV 2022poster

"Regarding Scene Graph Generation (SGG), coarse and fine predicates mix in the dataset due to the crowd-sourced labeling, and the long-tail problem is also pronounced. Given this tricky situation, many existing SGG methods treat the predicates equally and learn the model under the supervision of mix…

Cited by 31SourcePDFScholar
2022

MS2DG-Net: Progressive Correspondence Learning via Multiple Sparse Semantics Dynamic Graph

CVPR 2022poster

Establishing superior-quality correspondences in an image pair is pivotal to many subsequent computer vision tasks. Using Euclidean distance between correspondences to find neighbors and extract local information is a common strategy in previous works. However, most such works ignore similar sparse…

Cited by 67PDFcodeScholar
2022

RFNet: Unsupervised Network for Mutually Reinforcing Multi-Modal Image Registration and Fusion

CVPR 2022poster

In this paper, we propose a novel method to realize multi-modal image registration and fusion in a mutually reinforcing framework, termed as RFNet. We handle the registration in a coarse-to-fine fashion. For the first time, we exploit the feedback of image fusion to promote the registration accuracy…

Cited by 126PDFcodeScholar
2021

Appearance-based Loop Closure Detection via Bidirectional Manifold Representation Consensus

ICRA 2021poster

Loop closure detection (LCD), which aims to deal with the drift emerging when robots travel around the route, plays a key role in a simultaneous localization and mapping system. Unlike most current methods which focus on seeking an appropriate representation of images, we propose a novel two-stage p…

Cited by 7SourceScholar
2021

Motion Field Consensus with Locality Preservation: A Geometric Confirmation Strategy for Loop Closure Detection

IROS 2021poster

Loop closure detection (LCD), which aims to deal with the drift emerging when robots travel around the route, plays a key role in a simultaneous localization and mapping system. Unlike most current methods which focus on seeking an appropriate representation of images, we propose a novel two-stage p…

Cited by 2SourceScholar
2021

Robust Graph Autoencoder for Hyperspectral Anomaly Detection

ICASSP 2021accepted

Autoencoder can not only extract features in an unsupervised manner, but also selects samples out that differs significantly from others. However, autoencoder is sensitive to noise and anomalies during training, and the relationships between pixels are discarded. In order to tackle these problems, w…

Cited by 0SourceScholar
2021

T-Net: Effective Permutation-Equivariant Network for Two-View Correspondence Learning

ICCV 2021poster

We develop a conceptually simple, flexible, and effective framework (named T-Net) for two-view correspondence learning. Given a set of putative correspondences, we reject outliers and regress the relative pose encoded by the essential matrix, by an end-to-end framework, which is consisted of two nov…

Cited by 30PDFcodeScholar
2021

To Choose or to Fuse? Scale Selection for Crowd Counting

AAAI 2021technical

In this paper, we address the large scale variation problem in crowd counting by taking full advantage of the multi-scale feature representations in a multi-level network. We implement such an idea by keeping the counting error of a patch as small as possible with a proper feature level selection st…

2021

UTDN: An Unsupervised Two-Stream Dirichlet-Net for Hyperspectral Unmixing

ICASSP 2021accepted

Recently, the learning-based method has received much attention in the unsupervised hyperspectral unmixing, yet their ability to extract physically meaningful endmembers remains limited and the performance has not been satisfactory. In this paper, we propose a novel two-stream Dirichlet-net, termed…

Cited by 0SourceScholar
2021

Uniformity in Heterogeneity: Diving Deep Into Count Interval Partition for Crowd Counting

ICCV 2021poster

Recently, the problem of inaccurate learning targets in crowd counting draws increasing attention. Inspired by a few pioneering work, we solve this problem by trying to predict the indices of pre-defined interval bins of counts instead of the count values themselves. However, an inappropriate interv…

Cited by 49PDFcodeScholar
2021

Unsupervised Stacked Capsule Autoencoder for Hyperspectral Image Classification

ICASSP 2021accepted

Since CapsNet [1] shattered all previous records of algorithms for image recognition, the capsule's conception has attracted bright attention. It interprets an object by the geometrical arrangement of parts. We think it can be transferred to hyperspectral images. In a hyperspectral data cube, each p…

Cited by 0SourceScholar
2020

Geometric Estimation via Robust Subspace Recovery

ECCV 2020poster

Geometric estimation from image point correspondences is the core procedure of many 3D vision problems, which is prevalently accomplished by random sampling techniques. In this paper, we consider the problem from an optimization perspective, to exploit the intrinsic linear structure of point corresp…

2020

MTL-NAS: Task-Agnostic Neural Architecture Search Towards General-Purpose Multi-Task Learning

CVPR 2020poster

We propose to incorporate neural architecture search (NAS) into general-purpose multi-task learning (GP-MTL). Existing NAS methods typically define different search spaces according to different tasks. In order to adapt to different task combinations (i.e., task sets), we disentangle the GP-MTL netw…

Cited by 109PDFcodeScholar
2020

Multi-Scale Progressive Fusion Network for Single Image Deraining

CVPR 2020poster

Rain streaks in the air appear in various blurring degrees and resolutions due to different distances from their positions to the camera. Similar rain patterns are visible in a rain image as well as its multi-scale (or multi-resolution) versions, which makes it possible to exploit such complementary…

Cited by 844PDFcodeScholar
2019

NDDR-CNN: Layerwise Feature Fusing in Multi-Task CNNs by Neural Discriminative Dimensionality Reduction

CVPR 2019poster

In this paper, we propose a novel Convolutional Neural Network (CNN) structure for general-purpose multi-task learning (MTL), which enables automatic feature fusing at every layer from different tasks. This is in contrast with the most widely used MTL CNN structures which empirically or heuristicall…

Cited by 347PDFcodeScholar
2019

Progressive Fusion Video Super-Resolution Network via Exploiting Non-Local Spatio-Temporal Correlations

ICCV 2019oral

Most previous fusion strategies either fail to fully utilize temporal information or cost too much time, and how to effectively fuse temporal information from consecutive frames plays an important role in video super-resolution (SR). In this study, we propose a novel progressive fusion network for v…

Cited by 337PDFcodeScholar
2018

Visual Homing via Guided Locality Preserving Matching

ICRA 2018poster

This study proposes a simple yet surprisingly effective feature matching approach, termed as guided locality preserving matching (GLPM), for visual homing of panoramic images. The key idea of our approach is merely to preserve the neighborhood structures of potential true matches between two panoram…

Cited by 16SourceScholar