← Search

miaomiao Liu

34 accepted papers

2025

DCHM: Depth-Consistent Human Modeling for Multiview Detection

ICCV 2025poster

Multiview pedestrian detection typically involves two stages: human modeling and pedestrian localization. Human modeling represents pedestrians in 3D space by fusing multiview information, making its quality crucial for detection accuracy. However, existing methods often introduce noise and have low…

Cited by 0SourcePDFScholar
2025

Joint Optimization of Neural Radiance Fields and Continuous Camera Motion from a Monocular Video

CVPR 2025poster

Neural Radiance Fields (NeRF) has demonstrated its superior capability to represent 3D geometry but require accurately precomputed camera poses during training. To mitigate this requirement, existing methods jointly optimize camera poses and NeRF often relying on good pose initialisation or depth pr…

2025

Puzzles: Unbounded Video-Depth Augmentation for Scalable End-to-End 3D Reconstruction

NeurIPS 2025poster

Multi-view 3D reconstruction remains a core challenge in computer vision. Recent methods, such as DUSt3R and its successors, directly regress pointmaps from image pairs without relying on known scene geometry or camera parameters. However, the performance of these models is constrained by the divers…

Cited by 0SourceScholar
2024

HashPoint: Accelerated Point Searching and Sampling for Neural Rendering

CVPR 2024highlight

In this paper we address the problem of efficient point searching and sampling for volume neural rendering. Within this realm two typical approaches are employed: rasterization and ray tracing. The rasterization-based methods enable real-time rendering at the cost of increased memory and lower fidel…

2024

LDP: Language-driven Dual-Pixel Image Defocus Deblurring Network

CVPR 2024poster

Recovering sharp images from dual-pixel (DP) pairs with disparity-dependent blur is a challenging task. Existing blur map-based deblurring methods have demonstrated promising results. In this paper we propose to the best of our knowledge the first framework to introduce the contrastive language-imag…

Cited by 12SourcePDFScholar
2024

Mining Supervision for Dynamic Regions in Self-Supervised Monocular Depth Estimation

CVPR 2024poster

This paper focuses on self-supervised monocular depth estimation in dynamic scenes trained on monocular videos. Existing methods jointly estimate pixel-wise depth and motion relying mainly on an image reconstruction loss. Dynamic regions remain a critical challenge for these methods due to the inher…

2024

Neural SDF Flow for 3D Reconstruction of Dynamic Scenes

ICLR 2024poster

In this paper, we tackle the problem of 3D reconstruction of dynamic scenes from multi-view videos. Previous dynamic scene reconstruction works either attempt to model the motion of 3D points in space, which constrains them to handle a single articulated object or require depth maps as input. By con…

2023

K3DN: Disparity-Aware Kernel Estimation for Dual-Pixel Defocus Deblurring

CVPR 2023poster

The dual-pixel (DP) sensor captures a two-view image pair in a single snapshot by splitting each pixel in half. The disparity occurs in defocus blurred regions between the two views of the DP pair, while the in-focus sharp regions have zero disparity. This motivates us to propose a K3DN framework fo…

Cited by 12SourcePDFScholar
2023

Parametric Depth Based Feature Representation Learning for Object Detection and Segmentation in Bird's-Eye View

ICCV 2023poster

Recent vision-only perception models for autonomous driving achieved promising results by encoding multi-view image features into Bird's-Eye-View (BEV) space. A critical step and the main bottleneck of these methods is transforming image features into the BEV coordinate frame. This paper focuses on…

Cited by 10PDFcodeScholar
2022

EditVAE: Unsupervised Parts-Aware Controllable 3D Point Cloud Shape Generation

AAAI 2022technical

This paper tackles the problem of parts-aware point cloud generation. Unlike existing works which require the point cloud to be segmented into parts a priori, our parts-aware editing and generation are performed in an unsupervised manner. We achieve this with a simple modification of the Variational…

Cited by 38SourcePDFScholar
2022

Non-Parametric Depth Distribution Modelling Based Depth Inference for Multi-View Stereo

CVPR 2022poster

Recent cost volume pyramid based deep neural networks have unlocked the potential of efficiently leveraging high-resolution images for depth inference from multi-view stereo. In general, those approaches assume that the depth of each pixel follows a unimodal distribution. Boundary pixels usually fol…

Cited by 45PDFcodeScholar
2022

Spatially Invariant Unsupervised 3D Object-Centric Learning and Scene Decomposition

ECCV 2022poster

"We tackle the problem of object-centric learning on point clouds, which is crucial for high-level relational reasoning and scalable machine intelligence. In particular, we introduce a framework, SPAIR3D, to factorize a 3D point cloud into a spatial mixture model where each component corresponds to…

2022

Weakly-Supervised Action Transition Learning for Stochastic Human Motion Prediction

CVPR 2022oral

We introduce the task of action-driven stochastic human motion prediction, which aims to predict multiple plausible future motions given a sequence of action labels and a short motion history. This differs from existing works, which predict motions that either do not respect any specific action cate…

Cited by 41PDFcodeScholar
2021

Dual Pixel Exploration: Simultaneous Depth Estimation and Image Restoration

CVPR 2021poster

The dual-pixel (DP) hardware works by splitting each pixel in half and creating an image pair in a single snapshot. Several works estimate depth/inverse depth by treating the DP pair as a stereo pair. However, dual-pixel disparity only occurs in image regions with the defocus blur. The heavy defocus…

Cited by 43PDFScholar
2020

History Repeats Itself: Human Motion Prediction via Motion Attention

ECCV 2020poster

Human motion prediction aims to forecast future human poses given a past motion. Whether based on recurrent or feed-forward neural networks, existing methods fail to model the observation that human motion tends to repeat itself, even for complex sports actions and cooking activities. Here, we intro…

2019

Bringing a Blurry Frame Alive at High Frame-Rate With an Event Camera

CVPR 2019oral

Event-based cameras can measure intensity changes (called 'events') with microsecond accuracy under high-speed motion and challenging lighting conditions. With the active pixel sensor (APS), the event camera allows simultaneous output of the intensity frames. However, the output images are captured…

Cited by 312PDFScholar
2019

Phase-Only Image Based Kernel Estimation for Single Image Blind Deblurring

CVPR 2019poster

The image motion blurring process is generally modelled as the convolution of a blur kernel with a latent image. Therefore, the estimation of the blur kernel is essentially important for blind image deblurring. Unlike existing approaches which focus on approaching the problem by enforcing various pr…

Cited by 80PDFScholar
2017

Indoor Scene Parsing With Instance Segmentation, Semantic Labeling and Support Relationship Inference

CVPR 2017poster

Over the years, indoor scene parsing has attracted a growing interest in the computer vision community. Existing methods have typically focused on diverse subtasks of this challenging problem. In particular, while some of them aim at segmenting the image into regions, such as object or surface insta…

Cited by 39PDFScholar
2015

A Fixed Viewpoint Approach for Dense Reconstruction of Transparent Objects

CVPR 2015poster

This paper addresses the problem of reconstructing the surface shape of transparent objects. The difficulty of this problem originates from the viewpoint dependent appearance of a transparent object, which quickly makes reconstruction methods tailored for diffuse surfaces fail disgracefully. In this…

Cited by 47SourcePDFScholar
2015

Indoor Scene Structure Analysis for Single Image Depth Estimation

CVPR 2015poster

We tackle the problem of single image depth estimation, which, without additional knowledge, suffers from many ambiguities. Unlike previous approaches that only reason locally, we propose to exploit the global structure of the scene to estimate its depth. To this end, we introduce a hierarchical rep…

Cited by 143SourcePDFScholar