← Search

Huanjing Yue

19 accepted papers

2026

F^2HDR: Two-Stage HDR Video Reconstruction via Flow Adapter and Physical Motion Modeling

CVPR 2026

Reconstructing High Dynamic Range (HDR) videos from sequences of alternating-exposure Low Dynamic Range (LDR) frames remains highly challenging, especially under dynamic scenes where cross-exposure inconsistencies and complex motion make inter-frame alignment difficult, leading to ghosting and detai

Cited by 0SourcecodeScholar
2026

SpiralDiff: Spiral Diffusion with LoRA for RGB-to-RAW Conversion Across Cameras

CVPR 2026

RAW images preserve superior fidelity and rich scene information compared to RGB, making them essential for tasks in challenging imaging conditions. To alleviate the high cost of data collection, recent RGB-to-RAW conversion methods aim to synthesize RAW images from RGB. However, they overlook two k

Cited by 0SourcecodeScholar
2025

Incomplete Multi-modal Brain Tumor Segmentation via Learnable Sorting State Space Model

CVPR 2025poster

Brain tumor segmentation plays a crucial role in clinical diagnosis, yet the frequent unavailability of certain MRI modalities poses a significant challenge. In this paper, we introduce the Learnable Sorting State Space Model (LS3M), a novel framework designed to maximize the utilization of availabl…

Cited by 0SourcePDFScholar
2025

Learning Adaptive Lighting via Channel-Aware Guidance

ICML 2025poster

Learning lighting adaptation is a crucial step in achieving good visual perception and supporting downstream vision tasks. Current research often addresses individual light-related challenges, such as high dynamic range imaging and exposure correction, in isolation. However, we identify shared funda…

2025

Learning Differential Pyramid Representation for Tone Mapping

NeurIPS 2025poster

Existing tone mapping methods operate on downsampled inputs and rely on handcrafted pyramids to recover high-frequency details. Existing tone mapping methods operate on downsampled inputs and rely on handcrafted pyramids to recover high-frequency details. These designs typically fail to preserve fin…

Cited by 0SourcecodeScholar
2025

Zero-shot Video Restoration and Enhancement Using Pre-Trained Image Diffusion Model

AAAI 2025technical

Diffusion-based zero-shot image restoration and enhancement models have achieved great success in various tasks of image restoration and enhancement. However, directly applying them to video restoration and enhancement results in severe temporal flickering artifacts. In this paper, we propose the fi…

2024

AUFormer: Vision Transformers are Parameter-Efficient Facial Action Unit Detectors

ECCV 2024poster

"Facial Action Units (AU) is a vital concept in the realm of affective computing, and AU detection has always been a hot research topic. Existing methods suffer from overfitting issues due to the utilization of a large number of learnable parameters on scarce AU-annotated datasets or heavy reliance…

2024

Efficient Screen Content Image Compression via Superpixel-based Content Aggregation and Dynamic Feature Fusion

IJCAI 2024poster

This paper addresses the challenge of efficiently compressing screen content images (SCIs) – computer generated images with unique attributes such as large uniform regions, sharp edges, and limited color palettes, which pose difficulties for conventional compression algorithms. We propose a Superpix…

Cited by 0SourcePDFScholar
2024

KeDuSR: Real-World Dual-Lens Super-Resolution via Kernel-Free Matching

AAAI 2024technical

Dual-lens super-resolution (SR) is a practical scenario for reference (Ref) based SR by utilizing the telephoto image (Ref) to assist the super-resolution of the low-resolution wide-angle image (LR input). Different from general RefSR, the Ref in dual-lens SR only covers the overlapped field of view…

2024

TMFormer: Token Merging Transformer for Brain Tumor Segmentation with Missing Modalities

AAAI 2024technical

Numerous techniques excel in brain tumor segmentation using multi-modal magnetic resonance imaging (MRI) sequences, delivering exceptional results. However, the prevalent absence of modalities in clinical scenarios hampers performance. Current approaches frequently resort to zero maps as substitutes…

Cited by 5SourcePDFScholar
2024

Virtual Scanning: Unsupervised Non-line-of-sight Imaging from Irregularly Undersampled Transients

NeurIPS 2024poster

Non-line-of-sight (NLOS) imaging allows for seeing hidden scenes around corners through active sensing. Most previous algorithms for NLOS reconstruction require dense transients acquired through regular scans over a large relay surface, which limits their applicability in realistic scenarios with ir…

2023

Dec-Adapter: Exploring Efficient Decoder-Side Adapter for Bridging Screen Content and Natural Image Compression

ICCV 2023poster

Natural image compression has been greatly improved in the deep learning era. However, the compression performance will be heavily degraded if the pretrained encoder is directly applied on screen content image compression. Meanwhile, we observe that parameter-efficient trans-fer learning (PETL) meth…

Cited by 14PDFScholar
2023

Recaptured Raw Screen Image and Video Demoiréing via Channel and Spatial Modulations

NeurIPS 2023poster

Capturing screen contents by smartphone cameras has become a common way for information sharing. However, these images and videos are often degraded by moiré patterns, which are caused by frequency aliasing between the camera filter array and digital display grids. We observe that the moiré patterns…

2022

Real-RawVSR: Real-World Raw Video Super-Resolution with a Benchmark Dataset

ECCV 2022poster

"In recent years, real image super-resolution (SR) has achieved promising results due to the development of SR datasets and corresponding real SR methods. In contrast, the field of real video SR is lagging behind, especially for real raw videos. Considering the superiority of raw image SR over sRGB…

2022

Reference-Based Speech Enhancement via Feature Alignment and Fusion Network

AAAI 2022technical

Speech enhancement aims at recovering a clean speech from a noisy input, which can be classified into single speech enhancement and personalized speech enhancement. Personalized speech enhancement usually utilizes the speaker identity extracted from the noisy speech itself (or a clean reference spee…

2021

Implicit Transformer Network for Screen Content Image Continuous Super-Resolution

NeurIPS 2021poster

Nowadays, there is an explosive growth of screen contents due to the wide application of screen sharing, remote cooperation, and online education. To match the limited terminal bandwidth, high-resolution (HR) screen contents may be downsampled and compressed. At the receiver side, the super-resolu…

2021

Spatio-temporal Contrastive Domain Adaptation for Action Recognition

CVPR 2021poster

Unsupervised domain adaptation (UDA) for human action recognition is a practical and challenging problem. Compared with image-based UDA, video-based UDA is comprehensive to bridge the domain shift on both spatial representation and temporal dynamics. Most previous works focus on short-term modeling…

Cited by 87PDFScholar
2020

Supervised Raw Video Denoising With a Benchmark Dataset on Dynamic Scenes

CVPR 2020poster

In recent years, the supervised learning strategy for real noisy image denoising has been emerging and has achieved promising results. In contrast, realistic noise removal for raw noisy videos is rarely studied due to the lack of noisy-clean pairs for dynamic scenes. Clean video frames for dynamic s…

Cited by 142PDFcodeScholar
2018

Image Alignment via Multi-Model Geometric Fitting and Hierarchical Homography Estimation

ICASSP 2018accepted

It is challenging to achieve accurate alignment for building images containing multiple planes. We propose a multi-model geometric fitting and hierarchical homography estimation method to improve the alignment performance for building images. We first extract scale-invariant feature transform (SIFT)…

Cited by 0SourceScholar