← Search

Hujun Bao

114 accepted papers

2026

D-Prism: Differentiable Primitives for Structured Dynamic Modeling

CVPR 2026

Capturing both geometry and rigid motion for structured dynamic objects, like multi-part assemblies or jointed mechanisms, remains a key challenge. Existing dynamic methods, such as deformable meshes or 3DGS, rely on unstructured representations and fail to jointly model suitable geometry and articu

Cited by 0SourcecodeScholar
2026

DiffWind: Physics-Informed Differentiable Modeling of Wind-Driven Object Dynamics

ICLR 2026poster

Modeling wind-driven object dynamics from video observations is highly challenging due to the invisibility and spatio–temporal variability of wind, as well as the complex deformations of objects. We present DiffWind, a physics-informed differentiable framework that unifies wind–object interaction mo…

Cited by 0SourcecodeScholar
2026

GaussianZoom: Progressive Zoom-in Generative 3D Gaussian Splatting with Geometric and Semantic Guidance

CVPR 2026

We introduce GaussianZoom, a generative zoom-in 3D reconstruction system with an iterative progressive framework that combines geometry-consistent scene modeling and multi-scale semantic reasoning to enable high-fidelity extreme zoom-in rendering from low-resolution inputs. To achieve this, we devel

Cited by 0SourceScholar
2026

One-Shot Refiner: Boosting Feed-forward Novel View Synthesis via One-Step Diffusion

AAAI 2026technical

We present a novel framework for high-fidelity novel view synthesis (NVS) from sparse images, addressing key limitations in recent feed-forward 3D Gaussian Splatting (3DGS) methods built on Vision Transformer (ViT) backbones. While ViT-based pipelines offer strong geometric priors, they are often co

Cited by 0SourcePDFScholar
2026

PhysSkin: Real-Time and Generalizable Physics-Based Animation via Self-Supervised Neural Skinning

CVPR 2026

Achieving real-time physics-based animation that generalizes across diverse 3D shapes and discretizations remains a fundamental challenge. We introduce PhysSkin, a physics-informed framework that addresses this challenge. In the spirit of Linear Blend Skinning, we learn continuous skinning fields as

Cited by 0SourcecodeScholar
2026

StreamingTalker: Audio-driven 3D Facial Animation with Autoregressive Diffusion Model

AAAI 2026technical

This paper focuses on the task of speech-driven 3D facial animation, which aims to generate realistic and synchronized facial motions driven by speech inputs. Recent methods have employed audio-conditioned diffusion models for 3D facial animation, achieving impressive results in generating expressiv

Cited by 0SourcePDFScholar
2025

AccidentalGS: 3D Gaussian Splatting from Accidental Camera Motion

ICCV 2025poster

Neural 3D modeling and novel view synthesis with Neural Radiance Fields (NeRF) or 3D Gaussian Splatting (3DGS) typically requires the multi-view images with wide baselines and accurate camera poses as input. However, scenarios with accidental camera motions are rarely studied. In this paper, we prop…

Cited by 0SourcePDFScholar
2025

AtlasGS: Atlanta-world Guided Surface Reconstruction with Implicit Structured Gaussians

NeurIPS 2025poster

3D reconstruction of indoor and urban environments is a prominent research topic with various downstream applications. However, existing geometric priors for addressing low-texture regions in indoor and urban settings often lack global consistency. Moreover, Gaussian Splatting and implicit SDF fiel…

Cited by 0SourcecodeScholar
2025

BlinkTrack: Feature Tracking over 80 FPS via Events and Images

ICCV 2025poster

Event cameras, known for their high temporal resolution and ability to capture asynchronous changes, have gained significant attention for their potential in feature tracking, especially in challenging conditions. However, event cameras lack the fine-grained texture information that conventional cam…

2025

Diffuman4D: 4D Consistent Human View Synthesis from Sparse-View Videos with Spatio-Temporal Diffusion Models

ICCV 2025poster

This paper addresses the challenge of high-fidelity view synthesis of humans with sparse-view videos as input. Previous methods solve the issue of insufficient observation by leveraging 4D diffusion models to generate videos at novel viewpoints. However, the generated videos from these models often…

2025

ETO+: Revisit the Refinement Stage in Efficient Feature Matching

IROS 2025

Recent feature matching approaches like ETO have focused on developing lightweight matching algorithms for real-time applications. However, their lack of cross-image feature interaction and sufficient refinement often lead to a decline in matching accuracy. To address these challenges, we propose ET

Cited by 0SourceScholar
2025

EgoAgent: A Joint Predictive Agent Model in Egocentric Worlds

ICCV 2025poster

Learning an agent model that behaves like humans--capable of jointly perceiving the environment, predicting the future, and taking actions from a first-person perspective--is a fundamental challenge in computer vision. Existing methods typically train separate models for these abilities, which fail…

2025

EnvGS: Modeling View-Dependent Appearance with Environment Gaussian

CVPR 2025poster

Reconstructing complex reflections in real-world scenes from 2D images is essential for achieving photorealistic novel view synthesis. Existing methods that utilize environment maps to model reflections from distant lighting often struggle with high-frequency reflection details and fail to account f…

2025

FreeTimeGS: Free Gaussian Primitives at Anytime Anywhere for Dynamic Scene Reconstruction

CVPR 2025poster

This paper addresses the challenge of reconstructing dynamic 3D scenes with complex motions. Some recent works define 3D Gaussian primitives in the canonical space and use deformation fields to map canonical primitives to observation spaces, achieving real-time dynamic view synthesis. However, these…

Cited by 0SourcePDFScholar
2025

GURecon: Learning Detailed 3D Geometric Uncertainties for Neural Surface Reconstruction

AAAI 2025technical

Neural surface representation has demonstrated remarkable success in the areas of novel view synthesis and 3D reconstruction. However, assessing the geometric quality of 3D reconstructions in the absence of ground truth mesh remains a significant challenge, due to its rendering-based optimization pr…

Cited by 0SourcePDFScholar
2025

GaussianUpdate: Continual 3D Gaussian Splatting Update for Changing Environments

ICCV 2025poster

Novel view synthesis with neural models has advanced rapidly in recent years, yet adapting these models to scene changes remains an open problem. Existing methods are either labor-intensive, requiring extensive model retraining, or fail to capture detailed types of changes over time. In this paper,…

Cited by 0SourcePDFScholar
2025

Hierarchy UGP: Hierarchy Unified Gaussian Primitive for Large-Scale Dynamic Scene Reconstruction

ICCV 2025poster

Recent advances in differentiable rendering have significantly improved dynamic street scene reconstruction. However, the complexity of large-scale scenarios and dynamic elements, such as vehicles and pedestrians, remains a substantial challenge. Existing methods often struggle to scale to large sce…

Cited by 0SourcePDFScholar
2025

InstaScene: Towards Complete 3D Instance Decomposition and Reconstruction from Cluttered Scenes

ICCV 2025poster

Humans can naturally identify and mentally complete occluded objects in cluttered environments. However, imparting similar cognitive ability to robotics remains challenging even with advanced reconstruction techniques, which models scenes as undifferentiated wholes and fails to recognize complete ob…

Cited by 0SourcePDFScholar
2025

IntrinsicControlNet: Cross-distribution Image Generation with Real and Unreal

ICCV 2025poster

Realistic images are usually produced by simulating light transportation results of 3D scenes using rendering engines. This framework can precisely control the output but is usually weak at producing photo-like images. Alternatively, diffusion models have seen great success in photorealistic image g…

Cited by 0SourcePDFScholar
2025

LightCity: An Urban Dataset for Outdoor Inverse Rendering and Reconstruction under Multi-illumination Conditions

ICCV 2025poster

We propose an outdoor scene dataset and propose a series of benchmarks based on it.Inverse rendering in urban scenes is pivotal for applications like autonomous driving and digital twins, yet it faces significant challenges due to complex illumination conditions, including multi-illumination and ind…

Cited by 0SourcePDFScholar
2025

LookCloser: Frequency-aware Radiance Field for Tiny-Detail Scene

CVPR 2025poster

Humans perceive and comprehend their surroundings through information spanning multiple frequencies. In immersive scenes, people naturally scan their environment to grasp its overall structure while examining fine details of objects that capture their attention. However, current NeRF frameworks prim…

Cited by 1SourcePDFScholar
2025

MoEE: Mixture of Emotion Experts for Audio-Driven Portrait Animation

CVPR 2025poster

The generation of talking avatars has achieved significant advancements in precise audio synchronization. However, crafting lifelike talking head videos requires capturing a broad spectrum of emotions and subtle facial expressions. Current methods face fundamental challenges: a) the absence of frame…

Cited by 1SourcePDFScholar
2025

Multi-view Reconstruction via SfM-guided Monocular Depth Estimation

CVPR 2025poster

This paper aims to reconstruct the scene geometry from multi-view images with strong robustness and high quality. Previous learning-based methods incorporate neural networks into the multi-view stereo matching and have shown impressive reconstruction results. However, due to the reliance on matching…

2025

ND-SDF: Learning Normal Deflection Fields for High-Fidelity Indoor Reconstruction

ICLR 2025spotlight

Neural implicit reconstruction via volume rendering has demonstrated its effectiveness in recovering dense 3D surfaces. However, it is non-trivial to simultaneously recover meticulous geometry and preserve smoothness across regions with differing characteristics. To address this issue, previous meth…

2025

Neuraloc: Visual Localization in Neural Implicit Map With Dual Complementary Features

ICRA 2025

Recently, neural radiance fields (NeRF) have gained significant attention in the field of visual localization. However, existing NeRF-based approaches either lack geometric constraints or require extensive storage for feature matching, limiting their practical applications. To address these challeng

Cited by 6SourcecodeScholar
2025

Precise Action-to-Video Generation Through Visual Action Prompts

ICCV 2025poster

We present visual action prompts, a unified action representation for action-to-video generation of complex high-DoF interactions while maintaining transferable visual dynamics across domains. Action-driven video generation faces a precision-generality tradeoff: existing methods using text, primitiv…

Cited by 0SourcePDFScholar
2025

Prompting Depth Anything for 4K Resolution Accurate Metric Depth Estimation

CVPR 2025poster

Prompts play a critical role in unleashing the power of language and vision foundation models for specific tasks. For the first time, we introduce prompting into depth foundation models, creating a new paradigm for metric depth estimation termed Prompt Depth Anything. Specifically, we use a low-cost…

2025

ReTracker: Exploring Image Matching for Robust Online Any Point Tracking

ICCV 2025poster

This paper aims to establish correspondences for a set of 2D query points across a video sequence in an online manner. Recent methods leverage future frames to achieve smooth point tracking at the current frame, but they still struggle to find points with significant viewpoint changes after long-ter…

Cited by 0SourcePDFScholar
2025

Ready-to-React: Online Reaction Policy for Two-Character Interaction Generation

ICLR 2025poster

This paper addresses the task of generating two-character online interactions. Previously, two main settings existed for two-character interaction generation: (1) generating one's motions based on the counterpart's complete motion sequence, and (2) jointly generating two-character motions based on s…

Cited by 0SourcePDFScholar
2025

SGFormer: Satellite-Ground Fusion for 3D Semantic Scene Completion

CVPR 2025poster

Recently, camera-based solutions have been extensively explored for scene semantic completion (SSC). Despite their success in visible areas, existing methods struggle to capture complete scene semantics due to frequent visual occlusions. To address this limitation, this paper presents the first sate…

2025

Scalable Multi-Session Visual SLAM in Large-Scale Scenes with Subgraph Optimization

ICRA 2025

Multi-session visual SLAM systems enable 6-DoF camera localization along with long-term maintenance and expansion of the global map, by utilizing image data from different sessions. However, in large-scale environments, these systems often suffer from severe scale drift. While modern SLAM systems at

Cited by 1SourceScholar
2025

SpatialCrafter: Unleashing the Imagination of Video Diffusion Models for Scene Reconstruction from Limited Observations

ICCV 2025poster

Novel view synthesis (NVS) boosts immersive experiences in computer vision and graphics. Existing techniques, though progressed, rely on dense multi-view observations, restricting their application. This work takes on the challenge of reconstructing photorealistic 3D scenes from sparse or single-vie…

Cited by 0SourcePDFScholar
2025

SpatialTrackerV2: Advancing 3D Point Tracking with Explicit Camera Motion

ICCV 2025poster

We present SpatialTrackerV2, a feed-forward 3D point tracking method for monocular videos. Going beyond modular pipelines built on off-the-shelf components for 3D tracking, our approach unifies the intrinsic connections between point tracking, monocular depth, and camera pose estimation into a high-…

Cited by 0SourcePDFScholar
2025

StreetCrafter: Street View Synthesis with Controllable Video Diffusion Models

CVPR 2025poster

This paper aims to tackle the problem of photorealistic view synthesis from vehicle sensors data. Recent advancements in neural scene representation have achieved notable success in rendering high-quality autonomous driving scenes,but the performance significantly degrades as the viewpoint deviates…

Cited by 7SourcePDFScholar
2025

UniRestore3D: A Scalable Framework For General Shape Restoration

ICLR 2025poster

Shape restoration aims to recover intact 3D shapes from defective ones, such as those that are incomplete, noisy, and low-resolution. Previous works have achieved impressive results in shape restoration subtasks thanks to advanced generative models. While effective for specific shape defects, they a…

Cited by 0SourcePDFScholar
2025

UniVerse: Unleashing the Scene Prior of Video Diffusion Models for Robust Radiance Field Reconstruction

ICCV 2025poster

This paper tackles the challenge of robust reconstruction, i.e., the task of reconstructing a 3D scene from a set of inconsistent multi-view images. Some recent works have attempted to simultaneously remove image inconsistencies and perform reconstruction by integrating image degradation modeling in…

2024

"BlinkVision: A Benchmark for Optical Flow, Scene Flow and Point Tracking Estimation using RGB Frames and Events"

ECCV 2024poster

"Recent advances in event-based vision suggest that they complement traditional cameras by providing continuous observation without frame rate limitations and high dynamic range which are well-suited for correspondence tasks such as optical flow and point tracking. However, so far there is still a l…

Cited by 4SourcePDFScholar
2024

3D-SceneDreamer: Text-Driven 3D-Consistent Scene Generation

CVPR 2024poster

Text-driven 3D scene generation techniques have made rapid progress in recent years. Their success is mainly attributed to using existing generative models to iteratively perform image warping and inpainting to generate 3D scenes. However these methods heavily rely on the outputs of existing models…

Cited by 8SourcePDFScholar
2024

4K4D: Real-Time 4D View Synthesis at 4K Resolution

CVPR 2024poster

This paper targets high-fidelity and real-time view synthesis of dynamic 3D scenes at 4K resolution. Recent methods on dynamic view synthesis have shown impressive rendering quality. However their speed is still limited when rendering high-resolution images. To overcome this problem we propose 4K4D…

2024

A Global Depth-Range-Free Multi-View Stereo Transformer Network with Pose Embedding

NeurIPS 2024poster

In this paper, we propose a novel multi-view stereo (MVS) framework that gets rid of the depth range prior. Unlike recent prior-free MVS methods that work in a pair-wise manner, our method simultaneously considers all the source images. Specifically, we introduce a Multi-view Disparity Attention (MD…

Cited by 0SourcePDFScholar
2024

CG-SLAM: Efficient Dense RGB-D SLAM in a Consistent Uncertainty-aware 3D Gaussian Field

ECCV 2024poster

"Recently neural radiance fields (NeRF) have been widely exploited as 3D representations for dense simultaneous localization and mapping (SLAM). Despite their notable successes in surface modeling and novel view synthesis, existing NeRF-based methods are hindered by their computationally intensive a…

2024

Detector-Free Structure from Motion

CVPR 2024poster

We propose a structure-from-motion framework to recover accurate camera poses and point clouds from unordered images. Traditional SfM systems typically rely on the successful detection of repeatable keypoints across multiple views as the first step which is difficult for texture-poor scenes and poor…

2024

ETO:Efficient Transformer-based Local Feature Matching by Organizing Multiple Homography Hypotheses

NeurIPS 2024poster

We tackle the efficiency problem of learning local feature matching.Recent advancements have given rise to purely CNN-based and transformer-based approaches, each augmented with deep learning techniques. While CNN-based methods often excel in matching speed, transformer-based methods tend to provide…

Cited by 3SourcePDFScholar
2024

Error-aware Sampling in Adaptive Shells for Neural Surface Reconstruction

IJCAI 2024poster

Neural implicit surfaces with signed distance functions (SDFs) achieve superior quality in 3D geometry reconstruction. However, training SDFs is time-consuming because it requires a great number of samples to calculate accurate weight distributions and a considerable amount of samples sampled from t…

2024

From Satellite to Ground: Satellite Assisted Visual Localization with Cross-view Semantic Matching

ICRA 2024poster

One of the key challenges of visual Simultaneous Localization and Mapping (SLAM) in large-scale environments is how to effectively use global localization to correct the cumulative errors from long-term tracking. This challenge presents itself in two main aspects: first, the difficulty for robots in…

Cited by 1SourceScholar
2024

GeneAvatar: Generic Expression-Aware Volumetric Head Avatar Editing from a Single Image

CVPR 2024poster

Recently we have witnessed the explosive growth of various volumetric representations in modeling animatable head avatars. However due to the diversity of frameworks there is no practical method to support high-level applications like 3D head avatar editing across different representations. In this…

2024

Generating Human Motion in 3D Scenes from Text Descriptions

CVPR 2024poster

Generating human motions from textual descriptions has gained growing research interest due to its wide range of applications. However only a few works consider human-scene interactions together with text conditions which is crucial for visual and physical realism. This paper focuses on the task of…

2024

Omnidirectional Dense SLAM for Back-to-back Fisheye Cameras

ICRA 2024poster

We propose a real-time visual-inertial dense SLAM system that utilizes the online data streams from back-to-back dual fisheye cameras setup, providing 360◦ coverage of the environment. Firstly, we employ a sliding-window-based front-end to estimate real-time poses from the binocular fisheye images a…

Cited by 2SourceScholar
2024

PNeRFLoc: Visual Localization with Point-Based Neural Radiance Fields

AAAI 2024technical

Due to the ability to synthesize high-quality novel views, Neural Radiance Fields (NeRF) has been recently exploited to improve visual localization in a known environment. However, the existing methods mostly utilize NeRF for data augmentation to improve the regression model training, and their perf…

2024

Relightable and Animatable Neural Avatar from Sparse-View Video

CVPR 2024highlight

This paper tackles the problem of creating relightable and animatable neural avatars from sparse-view (or monocular) videos of dynamic humans under unknown illumination. Previous neural human reconstruction methods produce animatable avatars from sparse views using deformed Signed Distance Fields (S…

2023

AutoRecon: Automated 3D Object Discovery and Reconstruction

CVPR 2023highlight

A fully automated object reconstruction pipeline is crucial for digital content creation. While the area of 3D reconstruction has witnessed profound developments, the removal of background to obtain a clean object model still relies on different forms of manual labor, such as bounding box labeling,…

2023

BlinkFlow: A Dataset to Push the Limits of Event-Based Optical Flow Estimation

IROS 2023poster

Event cameras provide high temporal precision, low data rates, and high dynamic range visual perception, which are well-suited for optical flow estimation. While data-driven optical flow estimation has obtained great success in RGB cameras, its generalization performance is seriously hindered in eve…

Cited by 38SourcecodeScholar
2023

CF-Font: Content Fusion for Few-Shot Font Generation

CVPR 2023poster

Content and style disentanglement is an effective way to achieve few-shot font generation. It allows to transfer the style of the font image in a source domain to the style defined with a few reference images in a target domain. However, the content feature extracted using a representative font migh…

2023

CP-SLAM: Collaborative Neural Point-based SLAM System

NeurIPS 2023poster

This paper presents a collaborative implicit neural simultaneous localization and mapping (SLAM) system with RGB-D image sequences, which consists of complete front-end and back-end modules including odometry, loop detection, sub-map fusion, and global refinement. In order to enable all these module…

Cited by 29SourcePDFScholar
2023

Compact Neural Volumetric Video Representations with Dynamic Codebooks

NeurIPS 2023poster

This paper addresses the challenge of representing high-fidelity volumetric videos with low storage cost. Some recent feature grid-based methods have shown superior performance of fast learning implicit neural representations from input 2D images. However, such explicit representations easily lead t…

2023

DPS-Net: Deep Polarimetric Stereo Depth Estimation

ICCV 2023poster

Stereo depth estimation usually struggles to deal with textureless scenes for both traditional and learning-based methods due to the inherent dependence on image correspondence matching. In this paper, we propose a novel neural network, i.e., DPS-Net, to exploit both the prior geometric knowledge an…

Cited by 28PDFScholar
2023

Hierarchical Generation of Human-Object Interactions with Diffusion Probabilistic Models

ICCV 2023poster

This paper presents a novel approach to generating the 3D motion of a human interacting with a target object, with a focus on solving the challenge of synthesizing long-range and diverse motions, which could not be fulfilled by existing auto-regressive models or path planning-based methods. We propo…

Cited by 34PDFcodeScholar
2023

I2-SDF: Intrinsic Indoor Scene Reconstruction and Editing via Raytracing in Neural SDFs

CVPR 2023poster

In this work, we present I^2-SDF, a new method for intrinsic indoor scene reconstruction and editing using differentiable Monte Carlo raytracing on neural signed distance fields (SDFs). Our holistic neural SDF-based framework jointly recovers the underlying shapes, incident radiance and materials fr…

2023

IntrinsicNeRF: Learning Intrinsic Neural Radiance Fields for Editable Novel View Synthesis

ICCV 2023poster

Existing inverse rendering combined with neural rendering methods can only perform editable novel view synthesis on object-specific scenes, while we present intrinsic neural radiance fields, dubbed IntrinsicNeRF, which introduce intrinsic decomposition into the NeRF-based neural rendering method and…

Cited by 58PDFcodeScholar
2023

Learning Human Mesh Recovery in 3D Scenes

CVPR 2023poster

We present a novel method for recovering the absolute pose and shape of a human in a pre-scanned scene given a single image. Unlike previous methods that perform sceneaware mesh optimization, we propose to first estimate absolute position and dense scene contacts with a sparse 3D CNN, and later enha…

2023

Learning Neural Volumetric Representations of Dynamic Humans in Minutes

CVPR 2023poster

This paper addresses the challenge of efficiently reconstructing volumetric videos of dynamic humans from sparse multi-view videos. Some recent works represent a dynamic human as a canonical neural radiance field (NeRF) and a motion field, which are learned from input videos through differentiable r…

2023

Multi-Modal Neural Radiance Field for Monocular Dense SLAM with a Light-Weight ToF Sensor

ICCV 2023poster

Light-weight time-of-flight (ToF) depth sensors are compact and cost-efficient, and thus widely used on mobile devices for tasks such as autofocus and obstacle detection. However, due to the sparse and noisy depth measurements, these sensors have rarely been considered for dense geometry reconstruct…

Cited by 32PDFcodeScholar
2023

PATS: Patch Area Transportation With Subdivision for Local Feature Matching

CVPR 2023poster

Local feature matching aims at establishing sparse correspondences between a pair of images. Recently, detector-free methods present generally better performance but are not satisfactory in image pairs with large scale differences. In this paper, we propose Patch Area Transportation with Subdivision…

Cited by 42SourcePDFScholar
2023

Representing Volumetric Videos As Dynamic MLP Maps

CVPR 2023poster

This paper introduces a novel representation of volumetric videos for real-time view synthesis of dynamic scenes. Recent advances in neural scene representations demonstrate their remarkable capability to model and render complex static scenes, but extending them to represent dynamic scenes is not s…

2023

SINE: Semantic-Driven Image-Based NeRF Editing With Prior-Guided Editing Field

CVPR 2023poster

Despite the great success in 2D editing using user-friendly tools, such as Photoshop, semantic strokes, or even text prompts, similar capabilities in 3D areas are still limited, either relying on 3D modeling skills or allowing editing within only a few categories. In this paper, we present a novel s…

2023

Self-Distillation Hashing for Efficient Hamming Space Retrieval

ICASSP 2023accepted

Deep hashing-based approaches have become the optimal solutions for large-scale image retrieval task due to their high computational efficiency and low storage burden. Some methods leverage a large teacher network to improve the retrieval performance of the small student network through knowledge di…

Cited by 0SourceScholar
2022

Active Boundary Loss for Semantic Segmentation

AAAI 2022technical

This paper proposes a novel active boundary loss for semantic segmentation. It can progressively encourage the alignment between predicted boundaries and ground-truth boundaries during end-to-end training, which is not explicitly enforced in commonly used cross-entropy loss. Based on the predicted b…

2022

Crossview Mapping with Graph-based Geolocalization on City-Scale Street Maps

ICRA 2022poster

3D environment mapping has been actively stud-ied recently with the development of autonomous driving and augmented reality. Although many image-based methods are proposed due to their convenience and flexibility compared to other complex sensors, few works focus on fixing the inherent scale ambigui…

Cited by 5SourceScholar
2022

DELTAR: Depth Estimation from a Light-Weight ToF Sensor and RGB Image

ECCV 2022poster

"Light-weight time-of-flight (ToF) depth sensors are small, cheap, low-energy and have been massively deployed on mobile devices for the purposes like autofocus, obstacle detection, etc. However, due to their specific measurements (depth distribution in a region instead of the depth value at a certa…

2022

Geometry-aware Two-scale PIFu Representation for Human Reconstruction

NeurIPS 2022accept

Although PIFu-based 3D human reconstruction methods are popular, the quality of recovered details is still unsatisfactory. In a sparse (e.g., 3 RGBD sensors) capture setting, the depth noise is typically amplified in the PIFu representation, resulting in flat facial surfaces and geometry-fallible bo…

Cited by 17SourcePDFScholar
2022

NICE-SLAM: Neural Implicit Scalable Encoding for SLAM

CVPR 2022poster

Neural implicit representations have recently shown encouraging results in various domains, including promising progress in simultaneous localization and mapping (SLAM). Nevertheless, existing methods produce over-smoothed scene reconstructions and have difficulty scaling up to large scenes. These l…

Cited by 766PDFcodeScholar
2022

NeuMesh: Learning Disentangled Neural Mesh-Based Implicit Field for Geometry and Texture Editing

ECCV 2022poster

"Very recently neural implicit rendering techniques have been rapidly evolved and shown great advantages in novel view synthesis and 3D scene reconstruction. However, existing neural rendering methods for editing purposes offer limited functionality, e.g., rigid transformation, or not applicable for…

2022

Neural 3D Scene Reconstruction With the Manhattan-World Assumption

CVPR 2022oral

This paper addresses the challenge of reconstructing 3D indoor scenes from multi-view images. Many previous works have shown impressive reconstruction results on textured objects, but they still have difficulty in handling low-textured planar regions, which are common in indoor scenes. An approach t…

Cited by 188PDFcodeScholar
2022

OnePose++: Keypoint-Free One-Shot Object Pose Estimation without CAD Models

NeurIPS 2022accept

We propose a new method for object pose estimation without CAD models. The previous feature-matching-based method OnePose has shown promising results under a one-shot setting which eliminates the need for CAD models or object-specific training. However, OnePose relies on detecting repeatable image k…

2022

SelfRecon: Self Reconstruction Your Digital Avatar From Monocular Video

CVPR 2022oral

We propose SelfRecon, a clothed human body reconstruction method that combines implicit and explicit representations to recover space-time coherent geometries from a monocular self-rotating human video. Explicit methods require a predefined template mesh for a given sequence, while the template is h…

Cited by 184PDFcodeScholar
2022

TalkingFlow: Talking Facial Landmark Generation with Multi-Scale Normalizing Flow Network

ICASSP 2022accepted

Deterministic models dominate the field of talking facial land-mark generation by directly mapping speech signals to a certain lip-sync facial landmark sequence, which often suffer from regression to the mean face. In contrast, probability generative models are more beneficial to handle complex data…

Cited by 0SourceScholar
2022

TotalSelfScan: Learning Full-body Avatars from Self-Portrait Videos of Faces, Hands, and Bodies

NeurIPS 2022accept

Recent advances in implicit neural representations make it possible to reconstruct a human-body model from a monocular self-rotation video. While previous works present impressive results of human body reconstruction, the quality of reconstructed face and hands are relatively low. The main reason i…

2022

VIP-SLAM: An Efficient Tightly-Coupled RGB-D Visual Inertial Planar SLAM

ICRA 2022poster

In this paper, we propose a tightly-coupled SLAM system fused with RGB, Depth, IMU and structured plane information. Traditional sparse points based SLAM systems always maintain a mass of map points to model the environment. Huge number of map points bring us a high computational complexity, making…

Cited by 30SourceScholar
2021

AD-NeRF: Audio Driven Neural Radiance Fields for Talking Head Synthesis

ICCV 2021poster

Generating high-fidelity talking head video by fitting with the input audio sequence is a challenging problem that receives considerable attentions recently. In this paper, we address this problem with the aid of neural scene representation networks. Our method is completely different from existing…

Cited by 452PDFcodeScholar
2021

Animatable Neural Radiance Fields for Modeling Dynamic Human Bodies

ICCV 2021poster

This paper addresses the challenge of reconstructing an animatable human model from a multi-view video. Some recent works have proposed to decompose a non-rigidly deforming scene into a canonical neural radiance field and a set of deformation fields that map observation-space points to the canonical…

Cited by 514PDFcodeScholar
2021

Coxgraph: Multi-Robot Collaborative, Globally Consistent, Online Dense Reconstruction System

IROS 2021poster

Real-time dense reconstruction has been extensively studied for its wide applications in computer vision and robotics, meanwhile much effort has been made for the multi-robot system which plays an irreplaceable role in complicated but time-critical scenarios, e.g., search and rescue tasks. In this p…

Cited by 14SourceScholar
2021

DeepPanoContext: Panoramic 3D Scene Understanding With Holistic Scene Context Graph and Relation-Based Optimization

ICCV 2021poster

Panorama images have a much larger field-of-view thus naturally encode enriched scene context information compared to standard perspective images, which however is not well exploited in the previous scene understanding methods. In this paper, we propose a novel method for panoramic 3D scene understa…

Cited by 40PDFcodeScholar
2021

Graph-Based Asynchronous Event Processing for Rapid Object Recognition

ICCV 2021poster

Different from traditional video cameras, event cameras capture asynchronous events stream in which each event encodes pixel location, trigger time, and the polarity of the brightness changes. In this paper, we introduce a novel graph-based framework for event cameras, namely SlideGCN. Unlike some r…

Cited by 110PDFScholar
2021

Learning Object-Compositional Neural Radiance Field for Editable Scene Rendering

ICCV 2021poster

Implicit neural rendering techniques have shown promising results for novel view synthesis. However, existing methods usually encode the entire scene as a whole, which is generally not aware of the object identity and limits the ability to the high-level editing tasks such as moving or adding furnit…

Cited by 356PDFScholar
2021

LoFTR: Detector-Free Local Feature Matching With Transformers

CVPR 2021poster

We present a novel method for local image feature matching. Instead of performing image feature detection, description, and matching sequentially, we propose to first establish pixel-wise dense matches at a coarse level and later refine the good matches at a fine level. In contrast to dense methods…

Cited by 1526PDFcodeScholar
2021

Location-Aware Single Image Reflection Removal

ICCV 2021poster

This paper proposes a novel location-aware deep-learning-based single image reflection removal method. Our network has a reflection detection module to regress a probabilistic reflection confidence map, taking multi-scale Laplacian features as inputs. This probabilistic map tells if a region is refl…

Cited by 110PDFcodeScholar
2021

Neural Body: Implicit Neural Representations With Structured Latent Codes for Novel View Synthesis of Dynamic Humans

CVPR 2021poster

This paper addresses the challenge of novel view synthesis for a human performer from a very sparse set of camera views. Some recent works have shown that learning implicit neural representations of 3D scenes achieves remarkable view synthesis quality given dense input views. However, the representa…

Cited by 862PDFcodeScholar
2021

NeuralRecon: Real-Time Coherent 3D Reconstruction From Monocular Video

CVPR 2021poster

We present a novel framework named NeuralRecon for real-time 3D scene reconstruction from a monocular video. Unlike previous methods that estimate single-view depth maps separately on each key-frame and fuse them later, we propose to directly reconstruct local surfaces represented as sparse TSDF vol…

Cited by 358PDFcodeScholar
2021

Recurrent Multi-View Alignment Network for Unsupervised Surface Registration

CVPR 2021poster

Learning non-rigid registration in an end-to-end manner is challenging due to the inherent high degrees of freedom and the lack of labeled training data. In this paper, we resolve these two challenges simultaneously. First, we propose to represent the non-rigid transformation with a point-wise combi…

Cited by 55PDFcodeScholar
2021

StereoPIFu: Depth Aware Clothed Human Digitization via Stereo Vision

CVPR 2021poster

In this paper, we propose StereoPIFu, which integrates the geometric constraints of stereo vision with implicit function representation of PIFu, to recover the 3D shape of the clothed human from a pair of low-cost rectified images. First, we introduce the effective voxel-aligned features from a ster…

Cited by 82PDFScholar
2021

VS-Net: Voting With Segmentation for Visual Localization

CVPR 2021poster

Visual localization is of great importance in robotics and computer vision. Recently, scene coordinate regression based methods have shown good performance in visual localization in small static scenes. However, it still estimates camera poses from many inferior scene coordinates. To address this pr…

Cited by 57PDFcodeScholar
2021

You Don't Only Look Once: Constructing Spatial-Temporal Memory for Integrated 3D Object Detection and Tracking

ICCV 2021poster

Humans are able to continuously detect and track surrounding objects by constructing a spatial-temporal memory of the objects when looking around. In contrast, 3D object detectors in existing tracking-by-detection systems often search for objects in every new video frame from scratch, without fully…

Cited by 13PDFcodeScholar
2020

BCNet: Learning Body and Cloth Shape from A Single Image

ECCV 2020poster

In this paper, we consider the problem to automatically reconstruct garment and body shapes from a single near-front view RGB image. To this end, we propose a layered garment representation on top of SMPL and novelly make the skinning weight of garment independent of the body mesh, which significant…

2020

Disp R-CNN: Stereo 3D Object Detection via Shape Prior Guided Instance Disparity Estimation

CVPR 2020poster

In this paper, we propose a novel system named Disp R-CNN for 3D object detection from stereo images. Many recent works solve this problem by first recovering a point cloud with disparity estimation and then apply a 3D detector. The disparity map is computed for the entire image, which is costly and…

Cited by 145PDFcodeScholar
2020

Motion Capture from Internet Videos

ECCV 2020poster

Recent advances in image-based human pose estimation make it possible to capture 3D human motion from a single RGB video. However, the inherent depth ambiguity and self-occlusion in a single view prohibit the recovery of as high-quality motion as multi-view reconstruction. While multi-view videos ar…

2020

SMAP: Single-Shot Multi-Person Absolute 3D Pose Estimation

ECCV 2020poster

Recovering multi-person 3D poses with absolute scales from a single RGB image is a challenging problem due to the inherent depth and scale ambiguity from a single view. Addressing this ambiguity requires to aggregate various cues over the entire image, such as body sizes, scene layouts, and inter-pe…

Cited by 132SourcePDFScholar
2020

SelfVoxeLO: Self-supervised LiDAR Odometry with Voxel-based Deep Neural Networks

CoRL 2020

Recent learning-based LiDAR odometry methods have demonstrated their competitiveness. However, most methods still face two substantial challenges: 1) the 2D projection representation of LiDAR data cannot effectively encode 3D structures from the point clouds; 2) the needs for a large amount of label

2019

Depth Completion From Sparse LiDAR Data With Depth-Normal Constraints

ICCV 2019poster

Depth completion aims to recover dense depth maps from sparse depth measurements. It is of increasing importance for autonomous driving and draws increasing attention from the vision community. Most of the current competitive methods directly train a network to learn a mapping from sparse depth inpu…

Cited by 254PDFScholar
2019

Fast and Robust Multi-Person 3D Pose Estimation From Multiple Views

CVPR 2019poster

This paper addresses the problem of 3D pose estimation for multiple people in a few calibrated camera views. The main challenge of this problem is to find the cross-view correspondences among noisy and incomplete 2D pose predictions. Most previous methods address this challenge by directly reasoning…

Cited by 266PDFScholar
2019

GIFT: Learning Transformation-Invariant Dense Visual Descriptors via Group CNNs

NeurIPS 2019poster

Finding local correspondences between images with different viewpoints requires local descriptors that are robust against geometric transformations. An approach for transformation invariance is to integrate out the transformations by pooling the features extracted from transformed versions of an ima…

Cited by 109SourcePDFScholar
2019

PVNet: Pixel-Wise Voting Network for 6DoF Pose Estimation

CVPR 2019oral

This paper addresses the challenge of 6DoF pose estimation from a single RGB image under severe occlusion or truncation. Many recent works have shown that a two-stage approach, which first detects keypoints and then solves a Perspective-n-Point (PnP) problem for pose estimation, achieves remarkable…

Cited by 1355PDFcodeScholar
2019

Prior Guided Dropout for Robust Visual Localization in Dynamic Environments

ICCV 2019poster

Camera localization from monocular images has been a long-standing problem, but its robustness in dynamic environments is still not adequately addressed. Compared with classic geometric approaches, modern CNN-based methods (e.g. PoseNet) have manifested the reliability against illumination or viewpo…

Cited by 59PDFcodeScholar
2019

Rapid and Robust Monocular Visual-Inertial Initialization with Gravity Estimation via Vertical Edges

IROS 2019poster

Monocular visual-inertial tracking without good initialization easily fails due to its non-linear nature. Rapid and accurate metric initialization is crucial. In this paper, we propose a novel monocular visual-inertial initialization method which can initialize the IMU states, camera poses, and scal…

Cited by 13SourceScholar
2018

ICE-BA: Incremental, Consistent and Efficient Bundle Adjustment for Visual-Inertial SLAM

CVPR 2018poster

Modern visual-inertial SLAM (VI-SLAM) achieves higher accuracy and robustness than pure visual SLAM, thanks to the complementariness of visual features and inertial measurements. However, jointly using visual and inertial measurements to optimize SLAM objective functions is a problem of high computa…

2017

Robust stereo matching with surface normal prediction

ICRA 2017poster

Traditional stereo matching approaches generally have problems in handling textureless regions, strong occlusions and reflective regions that do not satisfy a Lambertian surface assumption. In this paper, we propose to combine the predicted surface normal by deep learning to overcome these inherent…

Cited by 23SourceScholar