← Search

Jingyu Yang

27 accepted papers

2026

Active Inference for Micro-Gesture Recognition: EFE-Guided Temporal Sampling and Adaptive Learning

CVPR 2026

Micro-gestures are subtle and transient movements triggered by unconscious neural and emotional activities, holding great potential for human-computer interaction and clinical monitoring. However, their low amplitude, short duration, and strong inter-subject variability make existing deep models pro

Cited by 0SourceScholar
2026

F^2HDR: Two-Stage HDR Video Reconstruction via Flow Adapter and Physical Motion Modeling

CVPR 2026

Reconstructing High Dynamic Range (HDR) videos from sequences of alternating-exposure Low Dynamic Range (LDR) frames remains highly challenging, especially under dynamic scenes where cross-exposure inconsistencies and complex motion make inter-frame alignment difficult, leading to ghosting and detai

Cited by 0SourcecodeScholar
2026

InterCoser: Interactive 3D Character Creation with Disentangled Fine-Grained Features

AAAI 2026technical

This paper aims to interactively generate and edit disentangled 3D characters based on precise user instructions. Existing methods generate and edit 3D characters via rough and simple editing guidance and entangled representations, making it difficult to achieve precise and comprehensive control ove

Cited by 0SourcePDFScholar
2026

SpiralDiff: Spiral Diffusion with LoRA for RGB-to-RAW Conversion Across Cameras

CVPR 2026

RAW images preserve superior fidelity and rich scene information compared to RGB, making them essential for tasks in challenging imaging conditions. To alleviate the high cost of data collection, recent RGB-to-RAW conversion methods aim to synthesize RAW images from RGB. However, they overlook two k

Cited by 0SourcecodeScholar
2025

Learning Adaptive Lighting via Channel-Aware Guidance

ICML 2025poster

Learning lighting adaptation is a crucial step in achieving good visual perception and supporting downstream vision tasks. Current research often addresses individual light-related challenges, such as high dynamic range imaging and exposure correction, in isolation. However, we identify shared funda…

2025

Learning Differential Pyramid Representation for Tone Mapping

NeurIPS 2025poster

Existing tone mapping methods operate on downsampled inputs and rely on handcrafted pyramids to recover high-frequency details. Existing tone mapping methods operate on downsampled inputs and rely on handcrafted pyramids to recover high-frequency details. These designs typically fail to preserve fin…

Cited by 0SourcecodeScholar
2025

Zero-shot Video Restoration and Enhancement Using Pre-Trained Image Diffusion Model

AAAI 2025technical

Diffusion-based zero-shot image restoration and enhancement models have achieved great success in various tasks of image restoration and enhancement. However, directly applying them to video restoration and enhancement results in severe temporal flickering artifacts. In this paper, we propose the fi…

2024

AUFormer: Vision Transformers are Parameter-Efficient Facial Action Unit Detectors

ECCV 2024poster

"Facial Action Units (AU) is a vital concept in the realm of affective computing, and AU detection has always been a hot research topic. Existing methods suffer from overfitting issues due to the utilization of a large number of learnable parameters on scarce AU-annotated datasets or heavy reliance…

2024

Efficient Screen Content Image Compression via Superpixel-based Content Aggregation and Dynamic Feature Fusion

IJCAI 2024poster

This paper addresses the challenge of efficiently compressing screen content images (SCIs) – computer generated images with unique attributes such as large uniform regions, sharp edges, and limited color palettes, which pose difficulties for conventional compression algorithms. We propose a Superpix…

Cited by 0SourcePDFScholar
2024

KeDuSR: Real-World Dual-Lens Super-Resolution via Kernel-Free Matching

AAAI 2024technical

Dual-lens super-resolution (SR) is a practical scenario for reference (Ref) based SR by utilizing the telephoto image (Ref) to assist the super-resolution of the low-resolution wide-angle image (LR input). Different from general RefSR, the Ref in dual-lens SR only covers the overlapped field of view…

2024

LPSNet: End-to-End Human Pose and Shape Estimation with Lensless Imaging

CVPR 2024poster

Human pose and shape (HPS) estimation with lensless imaging is not only beneficial to privacy protection but also can be used in covert surveillance scenarios due to the small size and simple structure of this device. However this task presents significant challenges due to the inherent ambiguity of…

Cited by 1SourcePDFScholar
2024

Virtual Scanning: Unsupervised Non-line-of-sight Imaging from Irregularly Undersampled Transients

NeurIPS 2024poster

Non-line-of-sight (NLOS) imaging allows for seeing hidden scenes around corners through active sensing. Most previous algorithms for NLOS reconstruction require dense transients acquired through regular scans over a large relay surface, which limits their applicability in realistic scenarios with ir…

2023

Dec-Adapter: Exploring Efficient Decoder-Side Adapter for Bridging Screen Content and Natural Image Compression

ICCV 2023poster

Natural image compression has been greatly improved in the deep learning era. However, the compression performance will be heavily degraded if the pretrained encoder is directly applied on screen content image compression. Meanwhile, we observe that parameter-efficient trans-fer learning (PETL) meth…

Cited by 14PDFScholar
2023

Learning Semantic-Aware Disentangled Representation for Flexible 3D Human Body Editing

CVPR 2023poster

3D human body representation learning has received increasing attention in recent years. However, existing works cannot flexibly, controllably and accurately represent human bodies, limited by coarse semantics and unsatisfactory representation capability, particularly in the absence of supervised da…

Cited by 8SourcePDFScholar
2023

NaviNeRF: NeRF-based 3D Representation Disentanglement by Latent Semantic Navigation

ICCV 2023poster

3D representation disentanglement aims to identify, decompose, and manipulate the underlying explanatory factors of 3D data, which helps AI fundamentally understand our 3D world. This task is currently under-explored and poses great challenges: (i) the 3D representations are complex and in general c…

Cited by 11PDFcodeScholar
2023

Recaptured Raw Screen Image and Video Demoiréing via Channel and Spatial Modulations

NeurIPS 2023poster

Capturing screen contents by smartphone cameras has become a common way for information sharing. However, these images and videos are often degraded by moiré patterns, which are caused by frequency aliasing between the camera filter array and digital display grids. We observe that the moiré patterns…

2022

FOF: Learning Fourier Occupancy Field for Monocular Real-time Human Reconstruction

NeurIPS 2022accept

The advent of deep learning has led to significant progress in monocular human reconstruction. However, existing representations, such as parametric models, voxel grids, meshes and implicit neural representations, have difficulties achieving high-quality results and real-time speed at the same time.…

Cited by 39SourcePDFScholar
2022

Real-RawVSR: Real-World Raw Video Super-Resolution with a Benchmark Dataset

ECCV 2022poster

"In recent years, real image super-resolution (SR) has achieved promising results due to the development of SR datasets and corresponding real SR methods. In contrast, the field of real video SR is lagging behind, especially for real raw videos. Considering the superiority of raw image SR over sRGB…

2022

Reference-Based Speech Enhancement via Feature Alignment and Fusion Network

AAAI 2022technical

Speech enhancement aims at recovering a clean speech from a noisy input, which can be classified into single speech enhancement and personalized speech enhancement. Personalized speech enhancement usually utilizes the speaker identity extracted from the noisy speech itself (or a clean reference spee…

2021

Implicit Transformer Network for Screen Content Image Continuous Super-Resolution

NeurIPS 2021poster

Nowadays, there is an explosive growth of screen contents due to the wide application of screen sharing, remote cooperation, and online education. To match the limited terminal bandwidth, high-resolution (HR) screen contents may be downsampled and compressed. At the receiver side, the super-resolu…

2021

Spatio-temporal Contrastive Domain Adaptation for Action Recognition

CVPR 2021poster

Unsupervised domain adaptation (UDA) for human action recognition is a practical and challenging problem. Compared with image-based UDA, video-based UDA is comprehensive to bridge the domain shift on both spatial representation and temporal dynamics. Most previous works focus on short-term modeling…

Cited by 87PDFScholar
2021

Target-targeted Domain Adaptation for Unsupervised Semantic Segmentation

ICRA 2021poster

Semantic segmentation has attracted increasing attention due to its important role in self-driving, and it is often realized by supervised learning with large number of well labeled maps. However, the labeled images are hard to be obtained in most circumstances, and the common way for unsupervised s…

Cited by 13SourceScholar
2020

Supervised Raw Video Denoising With a Benchmark Dataset on Dynamic Scenes

CVPR 2020poster

In recent years, the supervised learning strategy for real noisy image denoising has been emerging and has achieved promising results. In contrast, realistic noise removal for raw noisy videos is rarely studied due to the lack of noisy-clean pairs for dynamic scenes. Clean video frames for dynamic s…

Cited by 142PDFcodeScholar
2018

Image Alignment via Multi-Model Geometric Fitting and Hierarchical Homography Estimation

ICASSP 2018accepted

It is challenging to achieve accurate alignment for building images containing multiple planes. We propose a multi-model geometric fitting and hierarchical homography estimation method to improve the alignment performance for building images. We first extract scale-invariant feature transform (SIFT)…

Cited by 0SourceScholar
2018

Image-Based PM2.5 Estimation and its Application on Depth Estimation

ICASSP 2018accepted

Air pollution is still a big threat to human health particularly for developing countries. It is highly demanding to measure air quality with daily-used devices such as smartphones. On the other hand, it is difficult to estimate the scene depth under the foul weather using traditional vision-based m…

Cited by 0SourceScholar