← Search

Xinghao Ding

62 accepted papers

2026

Decompose and Attribute: Boosting Generalizable Open-Set Object Detection via Objectness Score

AAAI 2026technical

Open-set object detection (OSOD) aims to recognize known object categories while localizing previously unseen instances. However, real-world scenarios often involve co-occurring domain shifts and novel object categories. Existing OSOD methods typically overlook domain shifts, relying on source-train

Cited by 0SourcePDFScholar
2026

JarvisEvo: Towards a Self-Evolving Photo Editing Agent with Synergistic Editor-Evaluator Optimization

CVPR 2026

Agent-based editing models have substantially advanced interactive experiences, processing quality, and creative flexibility. However, two critical challenges persist: (1) instruction hallucination--text-only chain-of-thought (CoT) reasoning cannot fully prevent factual errors due to inherent inform

Cited by 0SourcecodeScholar
2026

MMMamba: A Versatile Cross-Modal in Context Fusion Framework for Pan-Sharpening and Zero-Shot Image Enhancement

AAAI 2026technical

Pan-sharpening aims to generate high-resolution multispectral (HRMS) images by integrating a high-resolution panchromatic (PAN) image with its corresponding low-resolution multispectral (MS) image. To achieve effective fusion, it is crucial to fully exploit the complementary information between the

Cited by 0SourcePDFScholar
2026

MobileFusion: Mobile-Friendly Infrared and Visible Image Fusion via Structural Re-parameterization

ICML 2026poster

Deep neural networks have recently advanced infrared and visible image fusion (IVIF), but most existing methods rely on sophisticated yet redundant designs, which hinder real-time deployment on mobile devices with limited compute and memory. In this paper, we present MobileFusion, an extremely light…

Cited by 0SourceScholar
2026

Self-supervised Multiplex Consensus Mamba for General Image Fusion

AAAI 2026technical

Image fusion integrates complementary information from different modalities to generate high-quality fused images, thereby enhancing downstream tasks such as object detection and semantic segmentation. Unlike task-specific techniques that primarily focus on consolidating inter-modal information, gen

Cited by 0SourcePDFScholar
2026

Thinking in Dynamics: How Multimodal Large Language Models Perceive, Track, and Reason Dynamics in Physical 4D World

CVPR 2026

Humans inhabit a physical 4D world, where spatial geometry and semantic content evolve over time, forming a dynamic reality. While current Multimodal Large Language Models (MLLMs) demonstrate strong capabilities in understanding static visual inputs, it remains unclear whether they can effectively "

Cited by 0SourcecodeScholar
2025

AGLLDiff: Guiding Diffusion Models Towards Unsupervised Training-free Real-world Low-light Image Enhancement

AAAI 2025technical

Existing low-light image enhancement (LIE) methods have achieved noteworthy success in solving synthetic distortions, yet they often fall short in practical applications. The limitations arise from two inherent challenges in real-world LIE: 1) the collection of distorted/clean image pairs is often i…

Cited by 7SourcePDFScholar
2025

ASGS: Single-Domain Generalizable Open-Set Object Detection via Adaptive Subgraph Searching

ICCV 2025poster

Albeit existing Single-Domain Generalized Object Detection (Single-DGOD) methods enable models to generalize to unseen domains, most assume that the training and testing data share the same label space. In real-world scenarios, unseen domains often introduce previously unknown objects, a challenge t…

Cited by 0SourcePDFScholar
2025

Accelerated Diffusion via High-Low Frequency Decomposition for Pan-Sharpening

AAAI 2025technical

Pan-sharpening aims to preserve the spectral information of the multi-spectral (MS) image while leveraging the high-frequency details from the guided high-resolution panchromatic (PAN) image to enhance its spatial resolution. The key challenge is how to preserve the spectral information from the MS…

Cited by 0SourcePDFScholar
2025

Bayesian Gaussian Process ODEs via Double Normalizing Flows

AISTATS 2025poster

Gaussian processes have been used to model the vector field of continuous dynamical systems, which are characterized by a probabilistic ordinary differential equation (GP-ODE). Bayesian inference for these models has been extensively studied and applied in tasks such as time series prediction. Howev…

Cited by 0SourceScholar
2025

DPLUT: Unsupervised Low-light Image Enhancement with Lookup Tables and Diffusion Priors

AAAI 2025technical

Low-light image enhancement (LIE) aims at precisely and efficiently recovering an image degraded in poor illumination environments. Recent advanced LIE techniques are using deep neural networks, which require lots of low-normal light image pairs, network parameters, and computational resources. As a…

Cited by 6SourcePDFScholar
2025

Demeaned Sparse: Efficient Anomaly Detection by Residual Estimate

ICML 2025poster

Frequency-domain image anomaly detection methods can substantially enhance anomaly detection performance, however, they still lack an interpretable theoretical framework to guarantee the effectiveness of the detection process. We propose a novel test to detect anomalies in structural image via a Dem…

Cited by 0SourcePDFScholar
2025

Dissecting Generalized Category Discovery: Multiplex Consensus under Self-Deconstruction

ICCV 2025poster

Human perceptual systems excel at inducing and recognizing objects across both known and novel categories, a capability far beyond current machine learning frameworks. While generalized category discovery (GCD) aims to bridge this gap, existing methods predominantly focus on optimizing objective fun…

2025

Dynamic Category Queries Transformer for Generalized Few-shot Semantic Segmentation

ICASSP 2025accepted

Few-shot segmentation (FSS) tackles data scarcity using multiple priors, but its simplicity limits handling base and novel classes with limited data access. Generalized few-shot semantic segmentation (GFSS) enhances model performance for base classes with abundant data, while novel classes have limi…

Cited by 0SourceScholar
2025

DynamicVerse: A Physically-Aware Multimodal Framework for 4D World Modeling

NeurIPS 2025poster

Understanding the dynamic physical world, characterized by its evolving 3D structure, real-world motion, and semantic content with textual descriptions, is crucial for human-agent interaction and enables embodied agents to perceive and act within real environments with human‑like capabilities. Howev…

Cited by 0SourceScholar
2025

Efficient Dataset Distillation through Low-Rank Space Sampling

ICASSP 2025accepted

Huge amount of data is the key of the success of deep learning, however, redundant information impairs the generalization ability of the model and increases the burden of calculation. Dataset Distillation (DD) compresses the original dataset into a smaller but representative subset for high-quality…

Cited by 0SourceScholar
2025

Efficient Infrared Image Super-Resolution Reconstruction via Guided Filter Coefficients Estimation with Parallax Attention Mechanism

ICASSP 2025accepted

Due to the spectral range mismatch between the images, building an efficient infrared (IR) image super-resolution algorithm suitable for embedded devices remains a significant challenge. Given that visible images possess more abundant high-frequency information compared to infrared images, we utiliz…

Cited by 0SourceScholar
2025

FRN: Fractal-Based Recursive Spectral Reconstruction Network

NeurIPS 2025poster

Generating hyperspectral images (HSIs) from RGB images through spectral reconstruction can significantly reduce the cost of HSI acquisition. In this paper, we propose a Fractal-Based Recursive Spectral Reconstruction Network (FRN), which differs from existing paradigms that attempt to directly integ…

Cited by 0SourcecodeScholar
2025

JarvisArt: Liberating Human Artistic Creativity via an Intelligent Photo Retouching Agent

NeurIPS 2025poster

Photo retouching has become integral to contemporary visual storytelling, enabling users to capture aesthetics and express creativity. While professional tools such as Adobe Lightroom offer powerful capabilities, they demand substantial expertise and manual effort. In contrast, existing AI-based sol…

Cited by 0SourceScholar
2025

JarvisIR: Elevating Autonomous Driving Perception with Intelligent Image Restoration

CVPR 2025poster

Vision-centric perception systems often struggle with unpredictable and coupled weather degradations in the wild. Current solutions are often limited, as they either depend on specific degradation priors or suffer from significant domain gaps. To enable robust and autonomous operation in real-world…

Cited by 2SourcePDFScholar
2025

Pan-LUT: Efficient Pan-sharpening via Learnable Look-Up Tables

NeurIPS 2025oral

Recently, deep learning-based pan-sharpening algorithms have achieved notable advancements over traditional methods. However, deep learning-based methods incur substantial computational overhead during inference, especially with large images. This excessive computational demand limits the applicabil…

Cited by 0SourceScholar
2025

Sp3ctralMamba: Physics-Driven Joint State Space Model for Hyperspectral Image Reconstruction

AAAI 2025technical

Hyperspectral image (HSI) reconstruction aims to restore the original 3D HSIs from the 2D hyperspectral snapshot compressive images (SCIs). The key to high-fidelity HSI reconstruction lies in designing refined spatial and spectral attention mechanisms, which are crucial for generating fine-grained r…

Cited by 0SourcePDFScholar
2025

Track Any Anomalous Object:A Granular Video Anomaly Detection Pipeline

CVPR 2025poster

Video anomaly detection (VAD) is crucial in scenarios such as surveillance and autonomous driving, where timely detection of unexpected activities is essential. Albeit existing methods have primarily focused on detecting anomalous objects in videos--either by identifying anomalous frames or objects-…

Cited by 0SourcePDFScholar
2024

Implicit Foreground-Guided Network for Anomaly Detection and Localization

ICASSP 2024accepted

Anomaly detection plays an essential role in large-scale industrial manufacturing. However, reconstruction-based anomaly detection methods, as one of the mainstream methods, are prone to incorrectly detecting background noise as anomalous regions. Therefore, inspired by multi-task learning, we propo…

Cited by 0SourceScholar
2024

Progressive High-Frequency Reconstruction for Pan-Sharpening with Implicit Neural Representation

AAAI 2024technical

Pan-sharpening aims to leverage the high-frequency signal of the panchromatic (PAN) image to enhance the resolution of its corresponding multi-spectral (MS) image. However, deep neural networks (DNNs) tend to prioritize learning the low-frequency components during the training process, which limits…

Cited by 11SourcePDFScholar
2024

Unsupervised Pan-Sharpening via Mutually Guided Detail Restoration

AAAI 2024technical

Pan-sharpening is a task that aims to super-resolve the low-resolution multispectral (LRMS) image with the guidance of a corresponding high-resolution panchromatic (PAN) image. The key challenge in pan-sharpening is to accurately modeling the relationship between the MS and PAN images. While supervi…

Cited by 3SourcePDFScholar
2023

Learning a Simple Low-Light Image Enhancer From Paired Low-Light Instances

CVPR 2023poster

Low-light Image Enhancement (LIE) aims at improving contrast and restoring details for images captured in low-light conditions. Most of the previous LIE algorithms adjust illumination using a single input image with several handcrafted priors. Those solutions, however, often fail in revealing image…

2023

Self-Supervised Image Denoising Using Implicit Deep Denoiser Prior

AAAI 2023technical

We devise a new regularization for denoising with self-supervised learning. The regularization uses a deep image prior learned by the network, rather than a traditional predefined prior. Specifically, we treat the output of the network as a ``prior'' that we again denoise after ``re-noising.'' The n…

Cited by 2SourcePDFScholar
2022

A Robust Object Segmentation Network for UnderWater Scenes

ICASSP 2022accepted

Underwater object segmentation is one of the key technologies in the fields of marine biology research and autonomous underwater vehicles. The challenges of underwater object segmentation originate from two aspects, 1) the complex underwater environment and 2) the camouflage characteristics of marin…

Cited by 0SourceScholar
2022

A Two-Stage Contrastive Learning Framework For Imbalanced Aerial Scene Recognition

ICASSP 2022accepted

In real-world scenarios, aerial image datasets are generally class imbalanced, where the majority classes have rich samples, while the minority classes only have a few samples. Such class imbalanced datasets bring great challenges to aerial scene recognition. In this paper, we explore a novel two-st…

Cited by 0SourceScholar
2022

Adaptive Variational Nonlinear Chirp Mode Decomposition

ICASSP 2022accepted

Variational nonlinear chirp mode decomposition (VNCMD) is a recently introduced method for nonlinear chirp signal decomposition that has aroused notable attention in various fields. One limiting aspect of the method is that its performance relies heavily on the setting of the bandwidth parameter. To…

Cited by 0SourceScholar
2022

Knowledge Condensation Distillation

ECCV 2022poster

"Knowledge Distillation (KD) transfers the knowledge from a high-capacity teacher network to strengthen a smaller student. Existing methods focus on excavating the knowledge hints and transferring the whole knowledge to the student. However, the knowledge redundancy arises since the knowledge shows…

2022

Uncertainty Inspired Underwater Image Enhancement

ECCV 2022poster

"A main challenge faced in the deep learning-based Underwater Image Enhancement (UIE) is that the ground truth high-quality image is unavailable. Most of the existing methods first generate approximate reference maps and then train an enhancement network with certainty. This kind of method fails to…

2022

Underwater Image Enhancement Via Learning Water Type Desensitized Representations

ICASSP 2022accepted

We present a novel underwater image enhancement method termed SCNet to improve the image quality meanwhile cope with the degradation diversity caused by the water. SCNet is based on normalization schemes across both spatial and channel dimensions with the key idea of learning water type desensitized…

Cited by 0SourceScholar
2022

Unsupervised Underwater Image Restoration: From a Homology Perspective

AAAI 2022technical

Underwater images suffer from degradation due to light scattering and absorption. It remains challenging to restore such degraded images using deep neural networks since real-world paired data is scarcely available while synthetic paired data cannot approximate real-world data perfectly. In this pap…

2022

Unsupervised and Untrained Underwater Image Restoration Based on Physical Image Formation Model

ICASSP 2022accepted

Underwater images suffer from degradation caused by light scattering and absorption. Training a deep neural network to restore underwater images is challenging due to the labor-intensive data collection and the lack of paired data. To this end, we propose an unsupervised and untrained underwater ima…

Cited by 0SourceScholar
2021

Dual Bipartite Graph Learning: A General Approach for Domain Adaptive Object Detection

ICCV 2021poster

Domain Adaptive Object Detection (DAOD) relieves the reliance on large-scale annotated data by transferring the knowledge learned from a labeled source domain to a new unlabeled target domain. Recent DAOD approaches resort to local feature alignment in virtue of domain adversarial training in conjun…

Cited by 68PDFScholar
2021

I3Net: Implicit Instance-Invariant Network for Adapting One-Stage Object Detectors

CVPR 2021poster

Recent works on two-stage cross-domain detection have widely explored the local feature patterns to achieve more accurate adaptation results. These methods heavily rely on the region proposal mechanisms and ROI-based instance-level features to design fine-grained feature alignment modules with respe…

Cited by 91PDFScholar
2021

Noise2Grad: Extract Image Noise to Denoise

IJCAI 2021poster

In many image denoising tasks, the difficulty of collecting noisy/clean image pairs limits the application of supervised CNNs. We consider such a case in which paired data and noise statistics are not accessible, but unpaired noisy and clean images are easy to collect. To form the necessary supervis…

Cited by 11SourcePDFScholar
2021

Rain Streak Removal via Dual Graph Convolutional Network

AAAI 2021technical

Deep convolutional neural networks (CNNs) have become dominant in the single image de-raining area. However, most deep CNNs-based de-raining methods are designed by stacking vanilla convolutional layers, which can only be used to model local relations. Therefore, long-range contextual information is…

Cited by 146SourcePDFScholar
2021

TRAR: Routing the Attention Spans in Transformer for Visual Question Answering

ICCV 2021poster

Due to the superior ability of global dependency modeling, Transformer and its variants have become the primary choice of many vision-and-language tasks. However, in tasks like Visual Question Answering (VQA) and Referring Expression Comprehension (REC), the multimodal prediction often requires visu…

Cited by 120PDFcodeScholar
2020

Harmonizing Transferability and Discriminability for Adapting Object Detectors

CVPR 2020poster

Recent advances in adaptive object detection have achieved compelling results in virtue of adversarial feature adaptation to mitigate the distributional shifts along the detection pipeline. Whilst adversarial adaptation significantly enhances the transferability of feature representations, the featu…

Cited by 361PDFcodeScholar
2019

JPEG Artifacts Reduction via Deep Convolutional Sparse Coding

ICCV 2019poster

To effectively reduce JPEG compression artifacts, we propose a deep convolutional sparse coding (DCSC) network architecture. We design our DCSC in the framework of classic learned iterative shrinkage-threshold algorithm. To focus on recognizing and separating artifacts only, we sparsely code the fea…

Cited by 139PDFScholar
2019

Look More Than Once: An Accurate Detector for Text of Arbitrary Shapes

CVPR 2019poster

Previous scene text detection methods have progressed substantially over the past years. However, limited by the receptive field of CNNs and the simple representations like rectangle bounding box or quadrangle adopted to describe text, previous methods may fall short when dealing with more challengi…

Cited by 328PDFScholar
2019

Lung Nodule Detection with a 3D ConvNet via IoU Self-normalization and Maxout Unit

ICASSP 2019accepted

The automatic pulmonary nodule detection in thoracic computed tomography (CT) scans plays a crucial role in the early diagnosis of lung cancer. In this paper, we propose a novel framework with a 3D convolutional network (ConvNet) for pulmonary nodule detection. To improve the efficiency and flexibil…

Cited by 0SourceScholar
2019

Progressive Feature Alignment for Unsupervised Domain Adaptation

CVPR 2019poster

Unsupervised domain adaptation (UDA) transfers knowledge from a label-rich source domain to a fully-unlabeled target domain. To tackle this task, recent approaches resort to discriminative domain transfer in virtue of pseudo-labels to enforce the class-level distribution alignment across the source…

Cited by 544PDFScholar
2019

Two-stream Multi-focus Image Fusion Based on the Latent Decision Map

ICASSP 2019accepted

The multi-focus image fusion with deep learning methods is mostly regarded as a two or three-category problem. Current systems utilize sliding windows to classify each pixel into focused or defocused, which is time consuming and requires post-processing such as denoising. In this paper, we propose a…

Cited by 0SourceScholar
2018

A Segmentation-aware Deep Fusion Network for Compressed Sensing MRI

ECCV 2018poster

Compressed sensing MRI is a classic inverse problem in the field of computational imaging, accelerating the MR imaging by measuring less k-space data. The deep neural network models provide the stronger representation ability and faster reconstruction compared with "shallow" optimization-based metho…

Cited by 34SourcePDFScholar
2018

Bindctnet: A Simple Binary Dct Network for Image Classification

ICASSP 2018accepted

Convolution neural networks play an important role in the image classification tasks. However, it is time consuming to train the network and the cost of memory resources is usually high. In this paper, a simple and effective network named BinDCTNet is presented by using the binary discrete cosine tr…

Cited by 0SourceScholar
2018

Man-Made Object Recognition from Underwater Optical Images Using Deep Learning and Transfer Learning

ICASSP 2018accepted

With the development of underwater optical sensors, manmade object recognition from underwater optical images has attracted wide attention. Deep learning methods have demonstrated impressive performance in object recognition tasks from natural images. However, it is difficult to collect large-scale…

Cited by 0SourceScholar
2017

Epithelium-stroma classification in histopathological images via convolutional neural networks and self-taught learning

ICASSP 2017accepted

Epithelium-stroma classification is always considered as an important preprocessing step for morphological quantitative analysis in image-based histological researches of oncologic diseases. However, large-scale accurate ground-truth labeling is expensive in histopathological image analysis, thus th…

Cited by 0SourceScholar
2017

PanNet: A Deep Network Architecture for Pan-Sharpening

ICCV 2017poster

We propose a deep network architecture for the pan-sharpening problem called PanNet. We incorporate domain-specific knowledge to design our PanNet architecture by focusing on the two aims of the pan-sharpening problem: spectral and spatial preservation. For spectral preservation, we add up-sampled m…

Cited by 800PDFScholar
2017

Removing Rain From Single Images via a Deep Detail Network

CVPR 2017poster

We propose a new deep network architecture for removing rain streaks from individual images based on the deep convolutional neural network (CNN). Inspired by the deep residual network (ResNet) that simplifies the learning process by changing the mapping form, we propose a deep detail network to dire…

Cited by 1377PDFScholar
2016

A Weighted Variational Model for Simultaneous Reflectance and Illumination Estimation

CVPR 2016poster

We propose a weighted variational model to estimate both the reflectance and the illumination from an observed image. We show that, though it is widely adopted for ease of modeling, the log-transformed image for this task is not ideal. Based on the previous investigation of the logarithmic transform…

Cited by 1183PDFScholar
2015

A novel pooling strategy for Full Reference Image Quality Assessment based on harmonic means

ICASSP 2015accepted

The most perceptual Full Reference Image Quality Assessment metrics (FR-IQA) shared a common two-step model; local quality measurement, and pooling. In this letter, a novel pooling strategy based on harmonic mean is proposed to predict the final quality score in FR-IQA. In contrast to arithmetic mea…

Cited by 0SourceScholar
2015

Fast magnetic susceptibility reconstruction using L0 norm of gradient

ICASSP 2015accepted

There is a growing interest in quantifying tissue susceptibility in MRI. However, the zeros in the dipole kernel makes the calculation of the magnetic susceptibility from the measured field to be an ill-posed problem. Recently, Bayesian regularization approaches have been utilized to enable accurate…

Cited by 0SourceScholar