← Search

Hongbo Fu

18 accepted papers

2026

Beyond VLM-Based Rewards: Diffusion-Native Latent Reward Modeling

ICML 2026poster

Preference optimization for diffusion models relies on reward functions that are both discriminative and computationally efficient. Vision-Language Models (VLMs) have emerged as powerful reward providers. However, their computation and memory cost can be substantial, and optimizing a latent diffusio…

Cited by 0SourceScholar
2025

GCRayDiffusion: Pose-Free Surface Reconstruction via Geometric Consistent Ray Diffusion

ICCV 2025poster

Accurate surface reconstruction from unposed images is crucial for efficient 3D object or scene creation. However, it remains challenging, particularly for the joint camera pose estimation. Previous approaches have achieved impressive pose-free surface reconstruction results in dense-view settings,…

2025

SketchVideo: Sketch-based Video Generation and Editing

CVPR 2025poster

Video generation and editing conditioned on text prompts or images have undergone significant advancements. However, challenges remain in accurately controlling global layout and geometry details solely by texts, and supporting motion control and local modification through images. In this paper, we…

Cited by 0SourcePDFScholar
2024

MonoHair: High-Fidelity Hair Modeling from a Monocular Video

CVPR 2024poster

Undoubtedly high-fidelity 3D hair is crucial for achieving realism artistic expression and immersion in computer graphics. While existing 3D hair modeling methods have achieved impressive performance the challenge of achieving high-quality hair reconstruction persists: they either require strict cap…

2024

Real-time 3D-aware Portrait Video Relighting

CVPR 2024highlight

Synthesizing realistic videos of talking faces under custom lighting conditions and viewing angles benefits various downstream applications like video conferencing. However most existing relighting methods are either time-consuming or unable to adjust the viewpoints. In this paper we present the fir…

2023

JR2Net: Joint Monocular 3D Face Reconstruction and Reenactment

AAAI 2023technical

Face reenactment and reconstruction benefit various applications in self-media, VR, etc. Recent face reenactment methods use 2D facial landmarks to implicitly retarget facial expressions and poses from driving videos to source images, while they suffer from pose and expression preservation issues fo…

Cited by 3SourcePDFScholar
2022

LiDAL: Inter-Frame Uncertainty Based Active Learning for 3D LiDAR Semantic Segmentation

ECCV 2022poster

"We propose LiDAL, a novel active learning method for 3D LiDAR semantic segmentation by exploiting inter-frame uncertainty among LiDAR frames. Our core idea is that a well-trained model should generate robust results irrespective of viewpoints for scene scanning and thus the inconsistencies in model…

2022

NeuralHDHair: Automatic High-Fidelity Hair Modeling From a Single Image Using Implicit Neural Representations

CVPR 2022poster

Undoubtedly, high-fidelity 3D hair plays an indispensable role in digital humans. However, existing monocular hair modeling methods are either tricky to deploy in digital systems (e.g., due to their dependence on complex user interactions or large databases) or can produce only a coarse geometry. In…

Cited by 38PDFScholar
2022

TransFusion: Robust LiDAR-Camera Fusion for 3D Object Detection With Transformers

CVPR 2022poster

LiDAR and camera are two important sensors for 3D object detection in autonomous driving. Despite the increasing popularity of sensor fusion in this field, the robustness against inferior image conditions, e.g., bad illumination and sensor misalignment, is under-explored. Existing fusion methods are…

Cited by 806PDFcodeScholar
2021

Autoregressive Stylized Motion Synthesis With Generative Flow

CVPR 2021poster

Motion style transfer is an important problem in many computer graphics and computer vision applications, including human animation, games, and robotics. Most existing deep learning methods for this problem are supervised and trained by registered motion pairs. In addition, these methods are often l…

Cited by 47PDFScholar
2021

Normalized Human Pose Features for Human Action Video Alignment

ICCV 2021poster

We present a novel approach for extracting human pose features from human action videos. The goal is to let the pose features capture only the poses of the action while being invariant to other factors, including video backgrounds, the video subject's anthropometric characteristics and viewpoints. S…

Cited by 17PDFScholar
2021

PointDSC: Robust Point Cloud Registration Using Deep Spatial Consistency

CVPR 2021poster

Removing outlier correspondences is one of the critical steps for successful feature-based point cloud registration. Despite the increasing popularity of introducing deep learning methods in this field, spatial consistency, which is essentially established by a Euclidean transformation between point…

Cited by 355PDFcodeScholar
2021

VMNet: Voxel-Mesh Network for Geodesic-Aware 3D Semantic Segmentation

ICCV 2021poster

In recent years, sparse voxel-based methods have become the state-of-the-arts for 3D semantic segmentation of indoor scenes, thanks to the powerful 3D CNNs. Nevertheless, being oblivious to the underlying geometry, voxel-based methods suffer from ambiguous features on spatially close objects and str…

Cited by 75PDFcodeScholar
2020

D3Feat: Joint Learning of Dense Detection and Description of 3D Local Features

CVPR 2020oral

A successful point cloud registration often lies on robust establishment of sparse matches through discriminative 3D local features. Despite the fast evolution of learning-based 3D feature descriptors, little attention has been drawn to the learning of 3D feature detectors, even less for a joint lea…

Cited by 528PDFcodeScholar
2020

End-to-End Learning Local Multi-View Descriptors for 3D Point Clouds

CVPR 2020poster

In this work, we propose an end-to-end framework to learn local multi-view descriptors for 3D point clouds. To adopt a similar multi-view representation, existing studies use hand-crafted viewpoints for rendering in a preprocessing stage, which is detached from the subsequent descriptor learning sta…

Cited by 144PDFScholar
2020

JSENet: Joint Semantic Segmentation and Edge Detection Network for 3D Point Clouds

ECCV 2020poster

Semantic segmentation and semantic edge detection can be seen as two dual problems with close relationships in computer vision. Despite the fast evolution of learning-based 3D semantic segmentation methods, little attention has been drawn to the learning of 3D semantic edge detectors, even less to a…

2020

Lidar-Monocular Visual Odometry using Point and Line Features

ICRA 2020poster

We introduce a novel lidar-monocular visual odometry approach using point and line features. Compared to previous point-only based lidar-visual odometry, our approach leverages more environment structure information by introducing both point and line features into pose estimation. We provide a robus…

Cited by 90SourceScholar