← Search

Zhiguo Cao

64 accepted papers

2026

BokehCrafter: Taming Video Diffusion Models for Controllable Bokeh Rendering

AAAI 2026technical

Bokeh is used in photography to emphasize the selected subject by smoothly blurring the out-of-focus region with appealing highlights. While recent advances have achieved impressive results in rendering realistic blur, existing frameworks typically rely on disparity maps and bokeh-relevant inputs (e

Cited by 0SourcePDFScholar
2026

BokehFlow: Depth-Free Controllable Bokeh Rendering via Flow Matching

AAAI 2026technical

Bokeh rendering simulates the shallow depth-of-field effect in photography, enhancing visual aesthetics and guiding viewer attention to regions of interest. Although recent approaches perform well, rendering controllable bokeh without additional depth inputs remains a significant challenge. Existing

Cited by 0SourcePDFScholar
2026

DEFANet: Dual-Path Edge-Target Collaboration with Frequency-Aware Enhancement for Infrared Small Target Detection

AAAI 2026technical

Infrared small target detection is challenging due to limited target size and low signal-to-noise ratio. Unlike common targets, infrared small targets contain a higher proportion of edge pixels and exhibit blurred boundaries due to diffraction and quantization artifacts, making boundaries uniquely v

Cited by 0SourcePDFScholar
2026

DeFB: Decomposed Feature Learning for Real-Time Multi-Person Eyeblink Detection in Untrimmed In-the-Wild Videos

AAAI 2026technical

Multi-person eyeblink detection in untrimmed in-the-wild videos is a recently emerged and challenging task. Due to its significant spatio-temporal fine-grained characteristics compared to general actions, we empirically find that general action detectors, though effective in general domains, struggl

Cited by 0SourcePDFScholar
2026

Identity-Preserving Image-to-Video Generation via Reward-Guided Optimization

CVPR 2026

Recent advances in image-to-video (I2V) generation have achieved remarkable progress in synthesizing high-quality, temporally coherent videos from static images. Among all the applications of I2V, human-centric video generation includes a large portion. However, existing I2V models encounter difficu

Cited by 0SourcecodeScholar
2026

JUMP-Hand: Learning Joint-wise Uncertainty to Gate Mixture of View Experts for Multi-View 3D Hand Reconstruction

CVPR 2026

We propose JUMP-Hand, a novel multi-view 3D hand reconstruction method that explicitly models probabilistic joint-wise uncertainty as a gating mechanism for multi-view fusion. Existing approaches usually rely on naive pooling or implicit attention, overlooking that each hand joint exhibits varying v

Cited by 0SourcecodeScholar
2026

Light-X: Generative 4D Video Rendering with Camera and Illumination Control

ICLR 2026poster

Recent advances in illumination control extend image-based methods to video, yet still facing a trade-off between lighting fidelity and temporal consistency. Moving beyond relighting, a key step toward generative modeling of real-world scenes is the joint control of camera trajectory and illuminatio…

Cited by 12SourcecodeScholar
2026

Semi-Supervised High Dynamic Range Image Reconstructing via Bi-Level Uncertain Area Masking

AAAI 2026technical

Reconstructing high dynamic range (HDR) images from low dynamic range (LDR) bursts plays an essential role in the computational photography. Impressive progress has been achieved by learning-based algorithms which require LDR-HDR image pairs. However, these pairs are hard to obtain, which motivates

Cited by 0SourcePDFScholar
2026

Through the Lens of Contrast: Self-Improving Visual Reasoning in VLMs

ICLR 2026oral

Reasoning has emerged as a key capability of large language models. In linguistic tasks, this capability can be enhanced by self-improving techniques that refine reasoning paths for subsequent fine-tuning. However, extending these language-based self-improving approaches to vision language models (V…

Cited by 0SourcecodeScholar
2025

CH3Depth: Efficient and Flexible Depth Foundation Model with Flow Matching

CVPR 2025highlight

Depth estimation is a fundamental task in 3D vision. An ideal depth estimation model is expected to embrace meticulous detail, temporal consistency, and high efficiency. Although existing foundation models can perform well in certain specific aspects, most of them fall short of fulfilling all the ab…

Cited by 0SourcePDFScholar
2025

DoF-Gaussian: Controllable Depth-of-Field for 3D Gaussian Splatting

CVPR 2025poster

Recent advances in 3D Gaussian Splatting (3D-GS) have shown remarkable success in representing 3D scenes and generating high-quality, novel views in real-time. However, 3D-GS and its variants assume that input images are captured based on pinhole imaging and are fully in focus. This assumption limit…

2025

Exploring Contextual Attribute Density in Referring Expression Counting

CVPR 2025poster

Referring expression counting (REC) algorithms are for more flexible and interactive counting ability across varied fine-grained text expressions. However, the requirement for fine-grained attribute understanding poses challenges for prior arts, as they struggle to accurately align attribute informa…

2025

Free4D: Tuning-free 4D Scene Generation with Spatial-Temporal Consistency

ICCV 2025poster

We present Free4D, a novel tuning-free framework for 4D scene generation from a single image. Existing methods either focus on object-level generation, making scene-level generation infeasible, or rely on large-scale multi-view video datasets for expensive training, with limited generalization abili…

2025

MuGS: Multi-Baseline Generalizable Gaussian Splatting Reconstruction

ICCV 2025poster

We present Multi-Baseline Gaussian Splatting (MuGS), a generalized feed-forward approach for novel view synthesis that effectively handles diverse baseline settings, including sparse input views with both small and large baselines. Specifically, we integrate features from Multi-View Stereo (MVS) and…

2025

PandaPose: 3D Human Pose Lifting from a Single Image via Propagating 2D Pose Prior to 3D Anchor Space

NeurIPS 2025poster

3D human pose lifting from a single RGB image is a challenging task in 3D vision. Existing methods typically establish a direct joint-to-joint mapping from 2D to 3D poses based on 2D features. This formulation suffers from two fundamental limitations: inevitable error propagation from input predicte…

Cited by 0SourceScholar
2025

SRefiner: Soft-Braid Attention for Multi-Agent Trajectory Refinement

ICCV 2025poster

Accurate prediction of multi-agent future trajectories is crucial for autonomous driving systems to make safe and efficient decisions. Trajectory refinement has emerged as a key strategy to enhance prediction accuracy. However, existing refinement methods often overlook the topological relationships…

2025

TacoDepth: Towards Efficient Radar-Camera Depth Estimation with One-stage Fusion

CVPR 2025award

Radar-Camera depth estimation aims to predict dense and accurate metric depth by fusing input images and Radar data. Model efficiency is crucial for this task in pursuit of real-time processing on autonomous vehicles and robotic platforms. However, due to the sparsity of Radar returns, the prevailin…

2025

WildAvatar: Learning In-the-wild 3D Avatars from the Web

CVPR 2025poster

Existing research on avatar creation is typically limited to laboratory datasets, which require high costs against scalability and exhibit insufficient representation of the real world. On the other hand, the web abounds with off-the-shelf real-world human videos, but these videos vary in quality an…

Cited by 2SourcePDFScholar
2024

3D Multi-frame Fusion for Video Stabilization

CVPR 2024poster

In this paper we present RStab a novel framework for video stabilization that integrates 3D multi-frame fusion through volume rendering. Departing from conventional methods we introduce a 3D multi-frame perspective to generate stabilized images addressing the challenge of full-frame generation while…

2024

CrossGLG: LLM Guides One-shot Skeleton-based 3D Action Recognition in a Cross-level Manner

ECCV 2024poster

"Most existing one-shot skeleton-based action recognition focuses on raw low-level information (, joint location), and may suffer from local information loss and low generalization ability. To alleviate these, we propose to leverage text description generated from large language models (LLM) that co…

Cited by 7SourcePDFScholar
2024

DyBluRF: Dynamic Neural Radiance Fields from Blurry Monocular Video

CVPR 2024poster

Recent advancements in dynamic neural radiance field methods have yielded remarkable outcomes. However these approaches rely on the assumption of sharp input images. When faced with motion blur existing dynamic NeRF methods often struggle to generate high-quality novel views. In this paper we propos…

Cited by 10SourcePDFScholar
2024

S-DyRF: Reference-Based Stylized Radiance Fields for Dynamic Scenes

CVPR 2024poster

Current 3D stylization methods often assume static scenes which violates the dynamic nature of our real world. To address this limitation we present S-DyRF a reference-based spatio-temporal stylization method for dynamic neural radiance fields. However stylizing dynamic 3D scenes is inherently chall…

Cited by 4SourcePDFScholar
2024

Self-Distilled Depth Refinement with Noisy Poisson Fusion

NeurIPS 2024poster

Depth refinement aims to infer high-resolution depth with fine-grained edges and details, refining low-resolution results of depth estimation models. The prevailing methods adopt tile-based manners by merging numerous patches, which lacks efficiency and produces inconsistency. Besides, prior arts su…

2024

Self-Supervised Class-Agnostic Motion Prediction with Spatial and Temporal Consistency Regularizations

CVPR 2024poster

The perception of motion behavior in a dynamic environment holds significant importance for autonomous driving systems wherein class-agnostic motion prediction methods directly predict the motion of the entire point cloud. While most existing methods rely on fully-supervised learning the manual labe…

2024

Semi-supervised Class-Agnostic Motion Prediction with Pseudo Label Regeneration and BEVMix

AAAI 2024technical

Class-agnostic motion prediction methods aim to comprehend motion within open-world scenarios, holding significance for autonomous driving systems. However, training a high-performance model in a fully-supervised manner always requires substantial amounts of manually annotated data, which can be bot…

2024

The All-Seeing Project: Towards Panoptic Visual Recognition and Understanding of the Open World

ICLR 2024poster

We present the All-Seeing (AS) project: a large-scale dataset and model for recognizing and understanding everything in the open world. Using a scalable data engine that incorporates human feedback and efficient models in the loop, we create a new dataset (AS-1B) with over 1.2 billion regions annota…

2024

Unifying Automatic and Interactive Matting with Pretrained ViTs

CVPR 2024poster

Automatic and interactive matting largely improve image matting by respectively alleviating the need for auxiliary input and enabling object selection. Due to different settings on whether prompts exist they either suffer from weakness in instance completeness or region details. Also when dealing wi…

2024

Vision Transformer Off-the-Shelf: A Surprising Baseline for Few-Shot Class-Agnostic Counting

AAAI 2024technical

Class-agnostic counting (CAC) aims to count objects of interest from a query image given few exemplars. This task is typically addressed by extracting the features of query image and exemplars respectively and then matching their feature similarity, leading to an extract-then-match paradigm. In this…

2023

A2J-Transformer: Anchor-to-Joint Transformer Network for 3D Interacting Hand Pose Estimation From a Single RGB Image

CVPR 2023poster

3D interacting hand pose estimation from a single RGB image is a challenging task, due to serious self-occlusion and inter-occlusion towards hands, confusing similar appearance patterns between 2 hands, ill-posed joint position mapping from 2D to 3D, etc.. To address these, we propose to extend A2J-…

2023

Constraining Depth Map Geometry for Multi-View Stereo: A Dual-Depth Approach with Saddle-shaped Depth Cells

ICCV 2023poster

Learning-based multi-view stereo (MVS) methods deal with predicting accurate depth maps to achieve an accurate and complete 3D representation. Despite the excellent performance, existing methods ignore the fact that a suitable depth geometry is also critical in MVS. In this paper, we demonstrate tha…

Cited by 17PDFcodeScholar
2023

Fast Full-frame Video Stabilization with Iterative Optimization

ICCV 2023poster

Video stabilization refers to the problem of transforming a shaky video into a visually pleasing one. The question of how to strike a good trade-off between visual quality and computational speed has remained one of the open challenges in video stabilization. Inspired by the analogy between wobbly f…

Cited by 12PDFcodeScholar
2023

Find Beauty in the Rare: Contrastive Composition Feature Clustering for Nontrivial Cropping Box Regression

AAAI 2023technical

Automatic image cropping algorithms aim to recompose images like human-being photographers by generating the cropping boxes with improved composition quality. Cropping box regression approaches learn the beauty of composition from annotated cropping boxes. However, the bias of annotations leads to q…

Cited by 6SourcePDFScholar
2023

Infusing Definiteness into Randomness: Rethinking Composition Styles for Deep Image Matting

AAAI 2023technical

We study the composition style in deep image matting, a notion that characterizes a data generation flow on how to exploit limited foregrounds and random backgrounds to form a training dataset. Prior art executes this flow in a completely random manner by simply going through the foreground pool or…

2023

Learning Second-Order Attentive Context for Efficient Correspondence Pruning

AAAI 2023technical

Correspondence pruning aims to search consistent correspondences (inliers) from a set of putative correspondences. It is challenging because of the disorganized spatial distribution of numerous outliers, especially when putative correspondences are largely dominated by outliers. It's more challengin…

2023

Matching Is Not Enough: A Two-Stage Framework for Category-Agnostic Pose Estimation

CVPR 2023highlight

Category-agnostic pose estimation (CAPE) aims to predict keypoints for arbitrary categories given support images with keypoint annotations. Existing approaches match the keypoints across the image for localization. However, such a one-stage matching paradigm shows inferior accuracy: the prediction h…

2023

Real-Time Multi-Person Eyeblink Detection in the Wild for Untrimmed Video

CVPR 2023poster

Real-time eyeblink detection in the wild can widely serve for fatigue detection, face anti-spoofing, emotion analysis, etc. The existing research efforts generally focus on single-person cases towards trimmed video. However, multi-person scenario within untrimmed videos is also important for practic…

2023

When Epipolar Constraint Meets Non-Local Operators in Multi-View Stereo

ICCV 2023poster

Learning-based multi-view stereo (MVS) method heavily relies on feature matching, which requires distinctive and descriptive representations. An effective solution is to apply non-local feature aggregation, e.g., Transformer. Albeit useful, these techniques introduce heavy computation overheads for…

Cited by 33PDFcodeScholar
2022

BokehMe: When Neural Rendering Meets Classical Rendering

CVPR 2022oral

We propose BokehMe, a hybrid bokeh rendering framework that marries a neural renderer with a classical physically motivated renderer. Given a single image and a potentially imperfect disparity map, BokehMe generates high-resolution photo-realistic bokeh effects with adjustable blur size, focal plane…

Cited by 48PDFcodeScholar
2022

C3P: Cross-Domain Pose Prior Propagation for Weakly Supervised 3D Human Pose Estimation

ECCV 2022poster

"This paper first proposes and solves weakly supervised 3D human pose estimation (HPE) problem in point cloud, via propagating the pose prior within unlabelled RGB-point cloud sequence to 3D domain. Our approach termed C3P does not require any labor-consuming 3D keypoint annotation for training. To…

2022

FADE: Fusing the Assets of Decoder and Encoder for Task-Agnostic Upsampling

ECCV 2022poster

"We consider the problem of task-agnostic feature upsampling in dense prediction where an upsampling operator is required to facilitate both region-sensitive tasks like semantic segmentation and detail-sensitive tasks such as image matting. Existing upsampling operators often can work well in either…

Cited by 48SourcePDFScholar
2022

MPIB: An MPI-Based Bokeh Rendering Framework for Realistic Partial Occlusion Effects

ECCV 2022poster

"Partial occlusion effects are a phenomenon that blurry objects near a camera are semi-transparent, resulting in partial appearance of occluded background. However, it is challenging for existing bokeh rendering methods to simulate realistic partial occlusion effects due to the missing information o…

2022

Represent, Compare, and Learn: A Similarity-Aware Framework for Class-Agnostic Counting

CVPR 2022poster

Class-agnostic counting (CAC) aims to count all instances in a query image given few exemplars. A standard pipeline is to extract visual features from exemplars and match them with query images to infer object counts. Two essential components in this pipeline are feature representation and similarit…

Cited by 109PDFcodeScholar
2022

Robust Object Detection with Inaccurate Bounding Boxes

ECCV 2022poster

"Learning accurate object detectors often requires large-scale training data with precise object bounding boxes. However, labeling such data is expensive and time-consuming. As the crowd-sourcing labeling process and the ambiguities of the objects may raise noisy bounding box annotations, the object…

2022

SAPA: Similarity-Aware Point Affiliation for Feature Upsampling

NeurIPS 2022accept

We introduce point affiliation into feature upsampling, a notion that describes the affiliation of each upsampled point to a semantic cluster formed by local decoder feature points with semantic similarity. By rethinking point affiliation, we present a generic formulation for generating upsampling k…

2021

TransView: Inside, Outside, and Across the Cropping View Boundaries

ICCV 2021poster

We show that relation modeling between visual elements matters in cropping view recommendation. Cropping view recommendation addresses the problem of image recomposition conditioned on the composition quality and the ranking of views (cropped sub-regions). This task is challenging because the visual…

Cited by 23PDFScholar
2020

3DV: 3D Dynamic Voxel for Action Recognition in Depth Video

CVPR 2020poster

For depth-based 3D action recognition, one essential issue is to represent 3D motion pattern effectively and efficiently. To this end, 3D dynamic voxel (3DV) is proposed as a novel 3D motion representation manner. With 3D space voxelization, the key idea of 3DV is to encode the 3D motion information…

Cited by 126PDFcodeScholar
2020

Measuring Generalisation to Unseen Viewpoints, Articulations, Shapes and Objects for 3D Hand Pose Estimation under Hand-Object Interaction

ECCV 2020poster

Articulations, Shapes and Objects for 3D Hand Pose Estimation under Hand-Object Interaction","We study how well different types of approaches generalise in the task of 3D hand pose estimation under single hand scenarios and hand-object interaction. We show that the accuracy of state-of-the-art metho…

2020

P2B: Point-to-Box Network for 3D Object Tracking in Point Clouds

CVPR 2020oral

Towards 3D object tracking in point clouds, a novel point-to-box network termed P2B is proposed in an end-to-end learning manner. Our main idea is to first localize potential target centers in 3D search area embedded with target information. Then point-driven 3D target proposal and verification are…

Cited by 199PDFcodeScholar
2020

Sparse-to-Dense Depth Completion Revisited: Sampling Strategy and Graph Construction

ECCV 2020poster

Depth completion is a widely studied problem of predicting a dense depth map from a sparse set of measurements and a single RGB image. In this work, we approach this problem by addressing two issues that have been under-researched in the open literature: sampling strategy (data term) and graph const…

Cited by 46SourcePDFScholar
2020

Structure-Guided Ranking Loss for Single Image Depth Prediction

CVPR 2020poster

Single image depth prediction is a challenging task due to its ill-posed nature and challenges with capturing ground truth for supervision. Large-scale disparity data generated from stereo photos and 3D videos is a promising source of supervision, however, such disparity data can only approximate th…

Cited by 214PDFcodeScholar
2020

Weighing Counts: Sequential Crowd Counting by Reinforcement Learning

ECCV 2020poster

We formulate counting as a sequential decision problem and present a novel crowd counting model solvable by deep reinforcement learning. In contrast to existing counting models that directly output count values, we divide one-step estimation into a sequence of much easier and more tractable sub-deci…

Cited by 97SourcePDFScholar
2019

A2J: Anchor-to-Joint Regression Network for 3D Articulated Pose Estimation From a Single Depth Image

ICCV 2019poster

For 3D hand and body pose estimation task in depth image, a novel anchor-based approach termed Anchor-to-Joint regression network (A2J) with the end-to-end learning ability is proposed. Within A2J, anchor points able to capture global-local spatial context information are densely set on depth image…

Cited by 221PDFcodeScholar
2019

From Open Set to Closed Set: Counting Objects by Spatial Divide-and-Conquer

ICCV 2019poster

Visual counting, a task that predicts the number of objects from an image/video, is an open-set problem by nature, i.e., the number of population can vary in [0,+[?]) in theory. However, the collected images and labeled count values are limited in reality, which means only a small closed set is obse…

Cited by 214PDFcodeScholar
2018

Monocular Relative Depth Perception With Web Stereo Data Supervision

CVPR 2018poster

In this paper we study the problem of monocular relative depth perception in the wild. We introduce a simple yet effective method to automatically generate dense relative depth annotations from web stereo images, and propose a new dataset that consists of diverse images as well as corresponding dens…

Cited by 253SourcePDFScholar
2017

When Unsupervised Domain Adaptation Meets Tensor Representations

ICCV 2017poster

Domain adaption (DA) allows machine learning methods trained on data sampled from one distribution to be applied to data sampled from another. It is thus of great practical importance to the application of such methods. Despite the fact that tensor representations are widely used in Computer Vision…

Cited by 89PDFcodeScholar