← Search

Shuaicheng Liu

68 accepted papers

2026

Action-Geometry Prediction with 3D Geometric Prior for Bimanual Manipulation

CVPR 2026

Bimanual manipulation requires policies that can reason about 3D geometry, anticipate how it evolves under action, and generate smooth, coordinated motions. However, existing methods typically rely on 2D features with limited spatial awareness, or require explicit point clouds that are difficult to

Cited by 0SourcecodeScholar
2026

Bridging RGB and RAW: Single-step Deterministic Flow with Homogeneous Aligned Guidance

ICML 2026poster

Reconstructing high-fidelity RAW sensor data from processed RGB images is a fundamental yet ill-posed problem, plagued by irreversible information loss and complex non-linear ISP transformations. While generative models offer high-quality reconstruction, they suffer from prohibitive computational co…

Cited by 0SourceScholar
2026

DMAligner: Enhancing Image Alignment via Diffusion Model Based View Synthesis

CVPR 2026

Image alignment is a fundamental task in computer vision with broad applications. Existing methods predominantly employ optical flow-based image warping. However, this technique is susceptible to common challenges such as occlusions and illumination variations, leading to degraded alignment visual q

Cited by 0SourcecodeScholar
2026

Efficient Hybrid SE(3)-Equivariant Visuomotor Flow Policy via Spherical Harmonics for Robot Manipulation

CVPR 2026

While existing equivariant methods enhance data efficiency, they suffer from high computational intensity, reliance on single-modality inputs, and instability when combined with fast-sampling methods. In this work, we propose E3Flow, a novel framework that addresses the critical limitations of equiv

Cited by 0SourcecodeScholar
2026

ExpoCM: Exposure-Aware One-Step Generative Single-Image HDR Reconstruction

CVPR 2026

Single-image HDR reconstruction aims to recover high dynamic range radiance from a single low dynamic range (LDR) input, but remains highly ill-posed due to detail saturation in over-exposed regions and noise amplification in under-exposed areas. While recent diffusion-based approaches offer powerfu

Cited by 0SourcecodeScholar
2026

HeRO: Hierarchical 3D Semantic Representation for Pose-Aware Object Manipulation

ICRA 2026poster

Imitation learning for robotic manipulation has progressed from 2D image policies to 3D representations that explicitly encode geometry. Yet purely geometric policies often lack explicit part-level semantics, which are critical for pose-aware manipulation (e.g., distinguishing a shoe's toe from heel…

2026

LaS-Comp: Zero-shot 3D Completion with Latent-Spatial Consistency

CVPR 2026

This paper introduces LaS-Comp, a zero-shot and category-agnostic approach that leverages the rich geometric priors of 3D foundation models to enable 3D shape completion across diverse types of partial observations. Our contributions are threefold: First, LaS-Comp harnesses these powerful generative

Cited by 0SourcecodeScholar
2026

RAW-Flow: Advancing RGB-to-RAW Image Reconstruction with Deterministic Latent Flow Matching

AAAI 2026technical

RGB-to-RAW reconstruction, or the reverse modeling of a camera Image Signal Processing (ISP) pipeline, aims to recover high-fidelity RAW data from RGB images. Despite notable progress, existing learning-based methods typically treat this task as a direct regression objective and still struggle with

Cited by 0SourcePDFScholar
2026

ZeroIDIR: Zero-Reference Illumination Degradation Image Restoration with Perturbed Consistency Diffusion Models

CVPR 2026

In this paper, we propose a zero-reference diffusion-based framework, named ZeroIDIR, for illumination degradation image restoration, which decouples the restoration process into adaptive illumination correction and diffusion-based reconstruction while being trained solely on low-quality degraded im

Cited by 0SourcecodeScholar
2025

Diff-Shadow: Global-guided Diffusion Model for Shadow Removal

AAAI 2025technical

We propose Diff-Shadow, a global-guided diffusion model for high-quality shadow removal. Previous transformer-based approaches can utilize global information to relate shadow and non-shadow regions but are limited in their synthesis ability and recover images with obvious boundaries. In contrast, di…

2025

Estimating 2D Camera Motion with Hybrid Motion Basis

ICCV 2025poster

Estimating 2D camera motion is a fundamental computer vision task that models the projection of 3D camera movements onto the 2D image plane. Current methods rely on either homography-based approaches, limited to planar scenes, or meshflow techniques that use grid-based local homographies but struggl…

2025

FlowPolicy: Enabling Fast and Robust 3D Flow-Based Policy via Consistency Flow Matching for Robot Manipulation

AAAI 2025technical

Robots can acquire complex manipulation skills by learning policies from expert demonstrations, which is often known as vision-based imitation learning. Generating policies based on diffusion and flow matching models has been shown to be effective, particularly in robotic manipulation tasks. However…

2025

HybridReg: Robust 3D Point Cloud Registration with Hybrid Motions

AAAI 2025technical

Scene-level point cloud registration is very challenging when considering dynamic foregrounds. Existing indoor datasets mostly assume rigid motions, so the trained models cannot robustly handle scenes with non-rigid motions. On the other hand, non-rigid datasets are mainly object-level, so the train…

2025

ISPDiffuser: Learning RAW-to-sRGB Mappings with Texture-Aware Diffusion Models and Histogram-Guided Color Consistency

AAAI 2025technical

RAW-to-sRGB mapping, or the simulation of the traditional camera image signal processor (ISP), aims to generate DSLR-quality sRGB images from raw data captured by smartphone sensors. Despite achieving comparable results to sophisticated handcrafted camera ISP solutions, existing learning-based metho…

2025

Learning Hazing to Dehazing: Towards Realistic Haze Generation for Real-World Image Dehazing

CVPR 2025poster

Existing real-world image dehazing methods primarily attempt to fine-tune pre-trained models or adapt their inference procedures, thus heavily relying on the pre-trained models and associated training data. Moreover, restoring heavily distorted information under dense haze requires generative diffus…

2025

Learning to See in the Extremely Dark

ICCV 2025poster

Learning-based methods have made promising advances in low-light RAW image enhancement, while their capability to extremely dark scenes where the environmental illuminance drops as low as 0.0001 lux remains to be explored due to the lack of corresponding datasets. To this end, we propose a paired-to…

2025

Realistic Noise Synthesis with Diffusion Models

AAAI 2025technical

Deep denoising models require extensive real-world training data, which is challenging to acquire. Current noise synthesis techniques struggle to accurately model complex noise distributions. We propose a novel Realistic Noise Synthesis Diffusor (RNSD) method using diffusion models to address these…

2025

Single Image Rolling Shutter Removal with Diffusion Models

AAAI 2025technical

We present RS-Diffusion, the first Diffusion Models-based method for single-frame Rolling Shutter (RS) correction. RS artifacts compromise visual quality of frames due to the row-wise exposure of CMOS sensors. Most previous methods have focused on multi-frame approaches, using temporal information f…

2025

The Parallel Pneumatic Artificial Muscle Platform Based on RBF Neural Network Compensation

IROS 2025

A two-degree-of-freedom parallel mechanism control system based on an adaptive learning rate and radial basis function (RBF) neural network controller is studied in this paper. The mechanism is composed of four pneumatic artificial muscles(PAM), forming two pairs of antagonistic single-degree-of-fre

Cited by 0SourceScholar
2025

Ultra High-Resolution Image Inpainting with Patch-Based Content Consistency Adapter

ICCV 2025poster

In this work, we present Patch-Adapter, an effective framework for high-resolution text-guided image inpainting. Unlike existing methods limited to lower resolutions, our approach achieves 4K+ resolution while maintaining precise content consistency and prompt alignment--two critical challenges in i…

2024

Efficient Meshflow and Optical Flow Estimation from Event Cameras

CVPR 2024poster

In this paper we explore the problem of event-based meshflow estimation a novel task that involves predicting a spatially smooth sparse motion field from event cameras. To start we generate a large-scale High-Resolution Event Meshflow (HREM) dataset which showcases its superiority by encompassing th…

2024

Eliminating Warping Shakes for Unsupervised Online Video Stitching

ECCV 2024poster

"In this paper, we retarget video stitching to an emerging issue, named warping shake, when extending image stitching to video stitching. It unveils the temporal instability of warped content in non-overlapping regions, despite image stitching having endeavored to preserve the natural structures. Th…

2024

FlowDiffuser: Advancing Optical Flow Estimation with Diffusion Models

CVPR 2024highlight

Optical flow estimation a process of predicting pixel-wise displacement between consecutive frames has commonly been approached as a regression task in the age of deep learning. Despite notable advancements this de facto paradigm unfortunately falls short in generalization performance when trained o…

2024

GLARE: Low Light Image Enhancement via Generative Latent Feature based Codebook Retrieval

ECCV 2024poster

"Most existing Low-light Image Enhancement (LLIE) methods either directly map Low-Light (LL) to Normal-Light (NL) images or use semantic or illumination maps as guides. However, the ill-posed nature of LLIE and the difficulty of semantic retrieval from impaired inputs limit these methods, especially…

2024

HandBooster: Boosting 3D Hand-Mesh Reconstruction by Conditional Synthesis and Sampling of Hand-Object Interactions

CVPR 2024poster

Reconstructing 3D hand mesh robustly from a single image is very challenging due to the lack of diversity in existing real-world datasets. While data synthesis helps relieve the issue the syn-to-real gap still hinders its usage. In this work we present HandBooster a new approach to uplift the data d…

2024

RecDiffusion: Rectangling for Image Stitching with Diffusion Models

CVPR 2024poster

Image stitching from different captures often results in non-rectangular boundaries which is often considered unappealing. To solve non-rectangular boundaries current solutions involve cropping which discards image content inpainting which can introduce unrelated content or warping which can distort…

2024

SpectralNeRF: Physically Based Spectral Rendering with Neural Radiance Field

AAAI 2024technical

In this paper, we propose SpectralNeRF, an end-to-end Neural Radiance Field (NeRF)-based architecture for high-quality physically based rendering from a novel spectral perspective. We modify the classical spectral rendering into two main steps, 1) the generation of a series of spectrum maps spanning…

2024

You Only Look Around: Learning Illumination-Invariant Feature for Low-light Object Detection

NeurIPS 2024poster

In this paper, we introduce YOLA, a novel framework for object detection in low-light scenarios. Unlike previous works, we propose to tackle this challenging problem from the perspective of feature learning. Specifically, we propose to learn illumination-invariant features through the Lambertian ima…

2023

AccFlow: Backward Accumulation for Long-Range Optical Flow

ICCV 2023poster

Recent deep learning-based optical flow estimators have exhibited impressive performance in generating local flows between consecutive frames. However, the estimation of long-range flows between distant frames, particularly under complex object deformation and large motion occlusion, remains a chall…

Cited by 24PDFcodeScholar
2023

Deep Homography Mixture for Single Image Rolling Shutter Correction

ICCV 2023poster

We present a deep homography mixture motion model for single image rolling shutter correction. Rolling shutter (RS) effects are often caused by row-wise exposure delay in the widely adopted CMOS sensor. Previous methods often require more than one frame for the correction, leading to data quality re…

Cited by 8PDFcodeScholar
2023

Explicit Motion Disentangling for Efficient Optical Flow Estimation

ICCV 2023poster

In this paper, we propose a novel framework for optical flow estimation that achieves a good balance between performance and efficiency. Our approach involves disentangling global motion learning from local flow estimation, treating global matching and local refinement as separate stages. We offer t…

Cited by 18PDFcodeScholar
2023

GAFlow: Incorporating Gaussian Attention into Optical Flow

ICCV 2023poster

Optical flow, or the estimation of motion fields from image sequences, is one of the fundamental problems in computer vision. Unlike most pixel-wise tasks that aim at achieving consistent representations of the same category, optical flow raises extra demands for obtaining local discrimination and s…

Cited by 32PDFcodeScholar
2023

Learning Optical Flow from Event Camera with Rendered Dataset

ICCV 2023poster

We study the problem of estimating optical flow from event cameras. One important issue is how to build a high-quality event-flow dataset with accurate event values and flow labels. Previous datasets are created by either capturing real scenes by event cameras or synthesizing from images with pasted…

Cited by 19PDFcodeScholar
2023

Low-Light Image Enhancement with Illumination-Aware Gamma Correction and Complete Image Modelling Network

ICCV 2023poster

This paper presents a novel network structure with illumination-aware gamma correction and complete image modelling to solve the low-light image enhancement problem. Low-light environments usually lead to less informative large-scale dark areas, directly learning deep representations from low-light…

Cited by 37PDFScholar
2023

MEFLUT: Unsupervised 1D Lookup Tables for Multi-exposure Image Fusion

ICCV 2023poster

In this paper, we introduce a new approach for high-quality multi-exposure image fusion (MEF). We show that the fusion weights of an exposure can be encoded into a 1D lookup table (LUT), which takes pixel intensity value as input and produces fusion weight as output. We learn one 1D LUT for each exp…

Cited by 19PDFcodeScholar
2023

SIRA-PCR: Sim-to-Real Adaptation for 3D Point Cloud Registration

ICCV 2023poster

Point cloud registration is essential for many applications. However, existing real datasets require extremely tedious and costly annotations, yet may not provide accurate camera poses. For the synthetic datasets, they are mainly object-level, so the trained models may not generalize well to real sc…

Cited by 22PDFcodeScholar
2023

Semi-supervised Deep Large-Baseline Homography Estimation with Progressive Equivalence Constraint

AAAI 2023technical

Homography estimation is erroneous in the case of large-baseline due to the low image overlay and limited receptive field. To address it, we propose a progressive estimation strategy by converting large-baseline homography into multiple intermediate ones, cumulatively multiplying these intermediate…

2023

Supervised Homography Learning with Realistic Dataset Generation

ICCV 2023poster

In this paper, we propose an iterative framework, which consists of two phases: a generation phase and a training phase, to generate realistic training data and yield a supervised homography network. In the generation phase, given an unlabeled image pair, we utilize the pre-estimated dominant plane…

Cited by 11PDFcodeScholar
2023

Uncertainty Guided Adaptive Warping for Robust and Efficient Stereo Matching

ICCV 2023poster

Correlation based stereo matching has achieved outstanding performance, which pursues cost volume between two feature maps. Unfortunately, current methods with a fixed trained model do not work uniformly well across various datasets, greatly limiting their real-world applicability. To tackle this is…

Cited by 24PDFScholar
2022

D2C-SR: A Divergence to Convergence Approach for Real-World Image Super-Resolution

ECCV 2022poster

"In this paper, we present D2C-SR, a novel framework for the task of real-world image super-resolution. As an ill-posed problem, the key challenge in super-resolution related tasks is there can be multiple predictions for a given low-resolution input. Most classical deep learning based approaches ig…

2022

Deep Constrained Least Squares for Blind Image Super-Resolution

CVPR 2022poster

In this paper, we tackle the problem of blind image super-resolution(SR) with a reformulated degradation model and two novel modules. Following the common practices of blind SR, our method proposes to improve both the kernel estimation as well as the kernel-based high-resolution image restoration. T…

Cited by 131PDFcodeScholar
2022

FINet: Dual Branches Feature Interaction for Partial-to-Partial Point Cloud Registration

AAAI 2022technical

Data association is important in the point cloud registration. In this work, we propose to solve the partial-to-partial registration from a new perspective, by introducing multi-level feature interactions between the source and the reference clouds at the feature extraction stage, such that the regi…

2022

Ghost-Free High Dynamic Range Imaging with Context-Aware Transformer

ECCV 2022poster

"High dynamic range (HDR) deghosting algorithms aim to generate ghost-free HDR images with realistic details. Restricted by the locality of the receptive field, existing CNN-based methods are typically prone to producing ghosting artifacts and intensity distortions in the presence of large motion an…

2022

Learning Optical Flow with Adaptive Graph Reasoning

AAAI 2022technical

Estimating per-pixel motion between video frames, known as optical flow, is a long-standing problem in video understanding and analysis. Most contemporary optical flow techniques largely focus on addressing the cross-image matching with feature similarity, with few methods considering how to explici…

2022

Practical Stereo Matching via Cascaded Recurrent Network With Adaptive Correlation

CVPR 2022oral

With the advent of convolutional neural networks, stereo matching algorithms have recently gained tremendous progress. However, it remains a great challenge to accurately extract disparities from real-world image pairs taken by consumer-level devices like smartphones, due to practical complicating f…

Cited by 318PDFcodeScholar
2022

RealFlow: EM-Based Realistic Optical Flow Dataset Generation from Videos

ECCV 2022poster

"Obtaining the ground truth labels from a video is challenging since the manual annotation of pixel-wise flow labels is prohibitively expensive and laborious. Besides, existing approaches try to adapt the trained model on synthetic datasets to authentic videos, which inevitably suffers from domain d…

2022

SceneSqueezer: Learning To Compress Scene for Camera Relocalization

CVPR 2022oral

Standard visual localization methods build a priori 3D model of a scene which is used to establish correspondences against the 2D keypoints in a query image. Storing these pre-built 3D scene models can be prohibitively expensive for large-scale environments, especially on mobile devices with limited…

Cited by 37PDFScholar
2022

Unsupervised Homography Estimation With Coplanarity-Aware GAN

CVPR 2022poster

Estimating homography from an image pair is a fundamental problem in image alignment. Unsupervised learning methods have received increasing attention in this field due to their promising performance and label-free training. However, existing methods do not explicitly consider the problem of plane i…

Cited by 54PDFcodeScholar
2021

DeepPanoContext: Panoramic 3D Scene Understanding With Holistic Scene Context Graph and Relation-Based Optimization

ICCV 2021poster

Panorama images have a much larger field-of-view thus naturally encode enriched scene context information compared to standard perspective images, which however is not well exploited in the previous scene understanding methods. In this paper, we propose a novel method for panoramic 3D scene understa…

Cited by 40PDFcodeScholar
2021

Holistic 3D Scene Understanding From a Single Image With Implicit Representation

CVPR 2021poster

We present a new pipeline for holistic 3D scene understanding from a single image, which could predict object shape, object pose and scene layout. As it is a highly ill-posed problem, existing methods usually suffer from inaccurate estimation of both shapes and layout especially for the cluttered sc…

Cited by 129PDFcodeScholar
2021

Motion Basis Learning for Unsupervised Deep Homography Estimation With Subspace Projection

ICCV 2021poster

In this paper, we introduce a new framework for unsupervised deep homography estimation. Our contributions are 3 folds. First, unlike previous methods that regress 4 offsets for a homography, we propose a homography flow representation, which can be estimated by a weighted sum of 8 pre-defined homog…

Cited by 70PDFcodeScholar
2021

NBNet: Noise Basis Learning for Image Denoising With Subspace Projection

CVPR 2021poster

In this paper, we introduce NBNet, a novel framework for image denoising. Unlike previous works, we propose to tackle this challenging problem from a new perspective: noise reduction by image-adaptive projection. Specifically, we propose to train a network that can separate signal and noise by learn…

Cited by 280PDFcodeScholar
2021

OMNet: Learning Overlapping Mask for Partial-to-Partial Point Cloud Registration

ICCV 2021poster

Point cloud registration is a key task in many computational fields. Previous correspondence matching based methods require the inputs to have distinctive geometric structures to fit a 3D rigid transformation according to point-wise sparse feature matches. However, the accuracy of transformation hea…

Cited by 207PDFcodeScholar
2021

Practical Wide-Angle Portraits Correction With Deep Structured Models

CVPR 2021poster

Wide-angle portraits often enjoy expanded views. However, they contain perspective distortions, especially noticeable when capturing group portrait photos, where the background is skewed and faces are stretched. This paper introduces the first deep learning based approach to remove such artifacts fr…

Cited by 23PDFcodeScholar
2021

UPFlow: Upsampling Pyramid for Unsupervised Optical Flow Learning

CVPR 2021poster

We present an unsupervised learning approach for optical flow estimation by improving the upsampling and learning of pyramid network. We design a self-guided upsample module to tackle the interpolation blur problem caused by bilinear upsampling between pyramid levels. Moreover, we propose a pyramid…

Cited by 111PDFcodeScholar
2020

Content-Aware Unsupervised Deep Homography Estimation

ECCV 2020poster

Homography estimation is a basic image alignment method in many applications. It is usually done by extracting and matching sparse feature points, which are error-prone in low-light and low-texture images. On the other hand, previous deep homography approaches use either synthetic images for supervi…

2020

DaST: Data-Free Substitute Training for Adversarial Attacks

CVPR 2020oral

Machine learning models are vulnerable to adversarial examples. For the black-box setting, current substitute attacks need pre-trained models to generate adversarial examples. However, pre-trained models are hard to obtain in real-world tasks. In this paper, we propose a data-free substitute trainin…

Cited by 208PDFcodeScholar
2019

DeepLiDAR: Deep Surface Normal Guided Depth Prediction for Outdoor Scene From Sparse LiDAR Data and Single Color Image

CVPR 2019poster

In this paper, we propose a deep learning architecture that produces accurate dense depth for the outdoor scene from a single color image and a sparse depth. Inspired by the indoor depth completion, our network estimates surface normals as the intermediate representation to produce dense depth, and…

Cited by 459PDFScholar
2017

Direct Photometric Alignment by Mesh Deformation

CVPR 2017poster

The choice of motion models is vital in applications like image/video stitching and video stabilization. Conventional methods explored different approaches ranging from simple global parametric models to complex per-pixel optical flow. Mesh-based warping methods achieve a good balance between comput…

Cited by 61PDFScholar