← Search

Kui Jiang

48 accepted papers

2026

Cross-ControlNet: Training-Free Fusion of Multiple Conditions for Text-to-Image Generation

ICLR 2026poster

Text-to-image diffusion models achieve impressive performance, but reconciling multiple spatial conditions usually requires costly retraining or labor intensive weight tuning. We introduce Cross-ControlNet, a training-free framework for text-to-image generation with multiple conditions. It exploits…

Cited by 0SourceScholar
2026

FedBPrompt: Federated Domain Generalization Person Re-Identification via Body Distribution Aware Visual Prompts

CVPR 2026

Federated Domain Generalization for Person Re-Identification (FedDG-ReID) aims to learn domain-invariant representations from decentralized data. Although Vision Transformers (ViTs) are widely adopted, their global attention often fails to distinguish pedestrians from high similarity backgrounds or

Cited by 0SourcecodeScholar
2026

ICLR: Inter-Chrominance and Luminance Interaction for Natural Color Restoration in Low-Light Image Enhancement

AAAI 2026technical

Low-Light Image Enhancement (LLIE) task aims at improving contrast while restoring details and textures for images captured in low-light conditions. HVI color space has made significant progress in this task by enabling precise decoupling of chrominance and luminance. However, for the interaction of

Cited by 0SourcePDFScholar
2026

Learning Depth from Past Selves: Self-Evolution Contrast for Robust Depth Estimation

AAAI 2026technical

Self-supervised depth estimation has gained significant attention in autonomous driving and robotics. However, existing methods exhibit substantial performance degradation under adverse weather conditions such as rain and fog, where reduced visibility critically impairs depth prediction. To address

Cited by 0SourcePDFScholar
2026

MambaOVSR: Multiscale Fusion with Global Motion Modeling for Chinese Opera Video Super-Resolution

AAAI 2026technical

Chinese opera is celebrated for preserving classical art. However, early filming equipment limitations have degraded videos of last-century performances by renowned artists (e.g., low frame rates and resolution), hindering archival efforts. Although space-time video super-resolution (STVSR) has adva

Cited by 0SourcePDFScholar
2026

Reliev3R: Relieving Feed-forward 3D Reconstruction from Multi-View Geometric Annotations

CVPR 2026

With recent advances, Feed-forward Reconstruction Models (FFRMs) have demonstrated great potential in reconstruction quality and adaptiveness to multiple downstream tasks. However, the excessive reliance on multi-view geometric annotations, e.g. 3D point maps and camera poses, makes the fully-superv

Cited by 0SourceScholar
2026

Semantics and Content Matter: Towards Multi-Prior Hierarchical Mamba for Image Deraining

AAAI 2026technical

Rain significantly degrades the performance of computer vision systems, particularly in applications like autonomous driving and video surveillance. While existing deraining methods have made considerable progress, they often struggle with fidelity of semantic and spatial details. To address these l

Cited by 0SourcePDFScholar
2026

Semi-LAR: Semi-supervised Contrastive Learning with Linear Attention for Removal of Nighttime Flares

ICML 2026poster

Lens flare removal is challenging due to the large spatial extent of flare artifacts and their entangle-ment with scene structures, while existing meth-ods heavily rely on large-scale paired data. We propose a semi-supervised flare removal frame-work that enables stable learning from unlabeled image…

Cited by 0SourceScholar
2025

Always Clear Depth: Robust Monocular Depth Estimation Under Adverse Weather

IJCAI 2025

Monocular depth estimation is critical for applications such as autonomous driving and scene reconstruction. While existing methods perform well under normal scenarios, their performance declines in adverse weather, due to challenging domain shifts and difficulties in extracting scene information. T

2025

Balancing Task-invariant Interaction and Task-specific Adaptation for Unified Image Fusion

ICCV 2025poster

Unified image fusion aims to integrate complementary information from multi-source images, enhancing image quality through a unified framework applicable to diverse fusion tasks. While treating all fusion tasks as a unified problem facilitates task-invariant knowledge sharing, it often overlooks tas…

2025

CLAIM: Mitigating Multilingual Object Hallucination in Large Vision-Language Models with Cross-Lingual Attention Intervention

ACL 2025long

Large Vision-Language Models (LVLMs) have demonstrated impressive multimodal abilities but remain prone to multilingual object hallucination, with a higher likelihood of generating responses inconsistent with the visual input when utilizing queries in non-English languages compared to English. Most…

Cited by 0SourcePDFScholar
2025

DashGaussian: Optimizing 3D Gaussian Splatting in 200 Seconds

CVPR 2025highlight

3D Gaussian Splatting (3DGS) renders pixels by rasterizing Gaussian primitives, where the rendering resolution and the primitive number, concluded as the optimization complexity, dominate the time cost in primitive optimization. In this paper, we propose DashGaussian, a scheduling scheme over the op…

Cited by 0SourcePDFScholar
2025

Debiased All-in-one Image Restoration with Task Uncertainty Regularization

AAAI 2025technical

All-in-one image restoration is a fundamental low-level vision task with significant real-world applications. The primary challenge lies in addressing diverse degradations within a single model. While current methods primarily exploit task prior information to guide the restoration models, they typi…

2025

Disentangle Nighttime Lens Flares: Self-supervised Generation-based Lens Flare Removal

AAAI 2025technical

Lens flares arise from light reflection and refraction within sensor arrays, whose diverse types include glow, veiling glare, reflective flare and so on. Existing methods are specialized for one specific type only, and overlook the simultaneous occurrence of multiple typed lens flares, which is comm…

Cited by 1SourcePDFScholar
2025

FUSE: Label-Free Image-Event Joint Monocular Depth Estimation via Frequency-Decoupled Alignment and Degradation-Robust Fusion

IROS 2025

Image-event joint depth estimation methods leverage complementary modalities for robust perception, yet face challenges in generalizability stemming from two factors: 1) limited annotated image-event-depth datasets causing insufficient cross-modal supervision, and 2) inherent frequency mismatches be

Cited by 0SourcecodeScholar
2025

Fast and Accurate Gigapixel Pathological Image Classification with Hierarchical Distillation Multi-Instance Learning

CVPR 2025poster

Although multi-instance learning (MIL) has succeeded in pathological image classification, it faces the challenge of high inference costs due to processing numerous patches from gigapixel whole slide images (WSIs).To address this, we propose HDMIL, a hierarchical distillation multi-instance learning…

2025

GMMamba: Group Masking Mamba for Whole Slide Image Classification

ICCV 2025poster

Recent advances in selective state space models (Mamba) have shown great promise in whole slide image (WSI) classification. Despite this, WSIs contain explicit local redundancy (similar patches) and irrelevant regions (uninformative instances), posing significant challenges for Mamba-based multi-ins…

Cited by 0SourcePDFScholar
2025

M3amba: Memory Mamba is All You Need for Whole Slide Image Classification

CVPR 2025poster

Multi-instance learning (MIL) has demonstrated impressive performance in whole slide image (WSI) analysis. However, existing approaches struggle with undesirable results and unbearable computational overhead due to the quadratic complexity of Transformers. Recently, Mamba has offered a feasible solu…

Cited by 0SourcePDFScholar
2025

OODML: Whole Slide Image Classification Meets Online Pseudo-Supervision and Dynamic Mutual Learning

AAAI 2025technical

Bag-label-based multi-instance learning (MIL) has demonstrated significant performance in whole slide image (WSI) analysis, particularly in pseudo-label-based learning schemes. However, due to inaccurate feature representation and interference, existing MIL methods often yield unreliable pseudo-labe…

Cited by 0SourcePDFScholar
2025

Reframing Gaussian Splatting Densification with Complexity-Density Consistency of Primitives

NeurIPS 2025poster

The essence of 3D Gaussian Splatting (3DGS) training is to smartly allocate Gaussian primitives, expressing complex regions with more primitives and vice versa. Prior researches typically mark out under-reconstructed regions in a rendering-loss-driven manner. However, such a loss-driven strategy i…

Cited by 0SourceScholar
2025

Spatial Annealing for Efficient Few-shot Neural Rendering

AAAI 2025technical

Neural Radiance Fields (NeRF) with hybrid representations have shown impressive capabilities for novel view synthesis, delivering high efficiency. Nonetheless, their performance significantly drops with sparse input views. Various regularization strategies have been devised to address these challeng…

2025

Spiking Meets Attention: Efficient Remote Sensing Image Super-Resolution with Attention Spiking Neural Networks

NeurIPS 2025poster

Spiking neural networks (SNNs) are emerging as a promising alternative to traditional artificial neural networks (ANNs), offering biological plausibility and energy efficiency. Despite these merits, SNNs are frequently hampered by limited capacity and insufficient representation power, yet remain un…

Cited by 0SourcecodeScholar
2025

The Parables of the Mustard Seed and the Yeast: Extremely Low-Budget, High-Performance Nighttime Semantic Segmentation

AAAI 2025technical

Nighttime Semantic Segmentation (NSS) is essential to many cutting-edge vision applications. However, existing technologies overly rely on massive labeled data, whose annotation is time-consuming and laborious. In this paper, we pioneer a new task focusing on exploring the potential of training stra…

Cited by 0SourcePDFScholar
2024

Dynamic Policy-Driven Adaptive Multi-Instance Learning for Whole Slide Image Classification

CVPR 2024highlight

Multi-Instance Learning (MIL) has shown impressive performance for histopathology whole slide image (WSI) analysis using bags or pseudo-bags. It involves instance sampling feature representation and decision-making. However existing MIL-based technologies at least suffer from one or more of the foll…

Cited by 6SourcePDFScholar
2024

FMRNet: Image Deraining via Frequency Mutual Revision

AAAI 2024technical

The wavelet transform has emerged as a powerful tool in deciphering structural information within images. And now, the latest research suggests that combining the prowess of wavelet transform with neural networks can lead to unparalleled image deraining results. By harnessing the strengths of both t…

2024

IQ-VFI: Implicit Quadratic Motion Estimation for Video Frame Interpolation

CVPR 2024poster

Advanced video frame interpolation (VFI) algorithms approximate intermediate motions between two input frames to synthesize intermediate frame. However they struggle to handle complex scenarios with curvilinear motions since they overlook the latent acceleration information between the input frames.…

Cited by 7SourcePDFScholar
2024

Improving Domain Generalization in Self-Supervised Monocular Depth Estimation via Stabilized Adversarial Training

ECCV 2024poster

"Learning a self-supervised Monocular Depth Estimation (MDE) model with great generalization remains significantly challenging. Despite the success of adversarial augmentation in the supervised learning generalization, naively incorporating it into self-supervised MDE models potentially causes over-…

Cited by 1SourcePDFScholar
2024

Learning a Spiking Neural Network for Efficient Image Deraining

IJCAI 2024poster

Recently, spiking neural networks (SNNs) have demonstrated substantial potential in computer vision tasks. In this paper, we present an Efficient Spiking Deraining Network, called ESDNet. Our work is motivated by the observation that rain pixel values will lead to a more pronounced intensity of spik…

2024

Learning from History: Task-agnostic Model Contrastive Learning for Image Restoration

AAAI 2024technical

Contrastive learning has emerged as a prevailing paradigm for high-level vision tasks, which, by introducing properly negative samples, has also been exploited for low-level vision tasks to achieve a compact optimization space to account for their ill-posed nature. However, existing methods rely on…

2024

Low-Light Face Super-resolution via Illumination, Structure, and Texture Associated Representation

AAAI 2024technical

Human face captured at night or in dimly lit environments has become a common practice, accompanied by complex low-light and low-resolution degradations. However, the existing face super-resolution (FSR) technologies and derived cascaded schemes are inadequate to recover credible textures. In this p…

2024

Mutuality Attribute Makes Better Video Anomaly Detection

ICASSP 2024accepted

Video anomaly detection (VAD) is an essential but challenging task. Existing prevalent methods focus on analyzing the reconstruction or prediction difference between normal and abnormal patterns through multiple deep features, e.g., optic flow. However, these approaches independently use deep featur…

Cited by 0SourceScholar
2024

OpticalDR: A Deep Optical Imaging Model for Privacy-Protective Depression Recognition

CVPR 2024poster

Depression Recognition (DR) poses a considerable challenge especially in the context of the growing concerns surrounding privacy. Traditional automatic diagnosis of DR technology necessitates the use of facial images undoubtedly expose the patient identity features and poses privacy risks. In order…

2023

From Generation to Suppression: Towards Effective Irregular Glow Removal for Nighttime Visibility Enhancement

IJCAI 2023poster

Most existing Low-Light Image Enhancement (LLIE) methods are primarily designed to improve brightness in dark regions, which suffer from severe degradation in nighttime images. However, these methods have limited exploration in another major visibility damage, the glow effects in real night scenes.…

Cited by 5SourcePDFScholar
2023

Refined Semantic Enhancement towards Frequency Diffusion for Video Captioning

AAAI 2023technical

Video captioning aims to generate natural language sentences that describe the given video accurately. Existing methods obtain favorable generation by exploring richer visual representations in encode phase or improving the decoding ability. However, the long-tailed problem hinders these attempts at…

2023

Store and Fetch Immediately: Everything Is All You Need for Space-Time Video Super-resolution

AAAI 2023technical

Existing space-time video super-resolution (ST-VSR) methods fail to achieve high-quality reconstruction since they fail to fully explore the spatial-temporal correlations, long-range components in particular. Although the recurrent structure for ST-VSR adopts bidirectional propagation to aggregate i…

2022

DANet: Image Deraining via Dynamic Association Learning

IJCAI 2022poster

Rain streaks and background components in a rainy input are highly correlated, making the deraining task a composition of the rain streak removal and background restoration. However, the correlation of these two components is barely considered, leading to unsatisfied deraining results. To this end,…

Cited by 21SourcePDFScholar
2022

Degrade Is Upgrade: Learning Degradation for Low-Light Image Enhancement

AAAI 2022technical

Low-light image enhancement aims to improve an image's visibility while keeping its visual naturalness. Different from existing methods, which tend to accomplish the relighting task directly, we investigate the intrinsic degradation and relight the low-light image while refining the details and colo…

2022

Rainy WCity: A Real Rainfall Dataset with Diverse Conditions for Semantic Driving Scene Understanding

IJCAI 2022poster

Scene understanding in adverse weather conditions (e.g. rainy and foggy days) has drawn increasing attention, arising some specific benchmarks and algorithms. However, scene segmentation under rainy weather is still challenging and under-explored due to the following limitations on the datasets and…

Cited by 32SourcePDFScholar
2022

Self-Supervised Learning on A Lightweight Low-Light Image Enhancement Model with Curve Refinement

ICASSP 2022accepted

Deep learning networks with deeper layers become a trend for their good performance but lacks the potential for real-time mobile deployment. Another challenge for paired training networks is the limited generalization capacity caused by the sample bias. To overcome these two challenges, we propose a…

Cited by 0SourceScholar
2022

Spatial-Temporal Space Hand-in-Hand: Spatial-Temporal Video Super-Resolution via Cycle-Projected Mutual Learning

CVPR 2022poster

Spatial-Temporal Video Super-Resolution (ST-VSR) aims to generate super-resolved videos with higher resolution (HR) and higher frame rate (HFR). Quite intuitively, pioneering two-stage based methods complete ST-VSR directly combining two sub-tasks: Spatial Video Super-Resolution (S-VSR) and Temporal…

Cited by 44PDFcodeScholar
2022

Unpaired Deep Image Deraining Using Dual Contrastive Learning

CVPR 2022poster

Learning single image deraining (SID) networks from an unpaired set of clean and rainy images is practical and valuable as acquiring paired real-world data is almost infeasible. However, without the paired data as the supervision, learning a SID network is challenging. Moreover, simply using existin…

Cited by 196PDFScholar
2022

VCD: View-Constraint Disentanglement for Action Recognition

ICASSP 2022accepted

Action recognition is a hot topic in computer vision due to its wide range of applications in urban surveillance. Although some methods are more advanced from an invariant view perspective, those approaches do not perform well for the viewpoint change. To address this issue, one possible solution is…

Cited by 0SourceScholar
2021

When Face Recognition Meets Occlusion: A New Benchmark

ICASSP 2021accepted

The existing face recognition datasets usually lack occlusion samples, which hinders the development of face recognition. Especially during the COVID-19 coronavirus epidemic, wearing a mask has become an effective means of preventing the virus spread. Traditional CNN-based face recognition models tr…

Cited by 0SourceScholar
2020

Attention-Guided Deraining Network Via Stage-Wise Learning

ICASSP 2020accepted

Due to diverse rain shapes, directions, densities as well as different distances to cameras, rain streaks in the air are interweaved and overlapped. However, most existing deraining methods are inherently oblivious this phenomenon and tend to learn a single rain streak layer to simulate this complex…

Cited by 0SourceScholar
2020

Multi-Scale Progressive Fusion Network for Single Image Deraining

CVPR 2020poster

Rain streaks in the air appear in various blurring degrees and resolutions due to different distances from their positions to the camera. Similar rain patterns are visible in a rain image as well as its multi-scale (or multi-resolution) versions, which makes it possible to exploit such complementary…

Cited by 844PDFcodeScholar
2019

Progressive Fusion Video Super-Resolution Network via Exploiting Non-Local Spatio-Temporal Correlations

ICCV 2019oral

Most previous fusion strategies either fail to fully utilize temporal information or cost too much time, and how to effectively fuse temporal information from consecutive frames plays an important role in video super-resolution (SR). In this study, we propose a novel progressive fusion network for v…

Cited by 337PDFcodeScholar