← Search

Yan Di

26 accepted papers

2026

PAMotion: Physics-Aware Motion Generation for Full-Body Interaction with Multiple Objects

CVPR 2026

We present PAMotion, a physics-aware diffusion framework for generating realistic full-body human interactions with multiple objects.Existing diffusion-based methods that jointly synthesize human and object motions often struggle to capture the intricate physical interactions--especially those invol

Cited by 0SourcecodeScholar
2025

SeaLion: Semantic Part-Aware Latent Point Diffusion Models for 3D Generation

CVPR 2025poster

Denoising diffusion probabilistic models have achieved significant success in point cloud generation, enabling numerous downstream applications, such as generative data augmentation and 3D model editing. However, little attention has been given to generating point clouds with point-wise segmentation…

Cited by 0SourcePDFScholar
2024

D-SCo: Dual-Stream Conditional Diffusion for Monocular Hand-Held Object Reconstruction

ECCV 2024poster

"Reconstructing hand-held objects from a single RGB image is a challenging task in computer vision. In contrast to prior works that utilize deterministic modeling paradigms, we employ a point cloud denoising diffusion model to account for the probabilistic nature of this problem. In the core, we int…

Cited by 2SourcePDFScholar
2024

EchoScene: Indoor Scene Generation via Information Echo over Scene Graph Diffusion

ECCV 2024poster

"We present EchoScene, an interactive and controllable generative model that generates 3D indoor scenes on scene graphs. EchoScene leverages a dual-branch diffusion model that dynamically adapts to scene graphs. Existing methods struggle to handle scene graphs due to varying numbers of nodes, multip…

Cited by 23SourcePDFScholar
2024

GeoGaussian: Geometry-aware Gaussian Splatting for Scene Rendering

ECCV 2024poster

"During the Gaussian Splatting optimization process, the scene geometry can gradually deteriorate if its structure is not deliberately preserved, especially in non-textured regions such as walls, ceilings, and furniture surfaces. This degradation significantly affects the rendering quality of novel…

Cited by 25SourcePDFScholar
2024

HiPose: Hierarchical Binary Surface Encoding and Correspondence Pruning for RGB-D 6DoF Object Pose Estimation

CVPR 2024poster

In this work we present a novel dense-correspondence method for 6DoF object pose estimation from a single RGB-D image. While many existing data-driven methods achieve impressive performance they tend to be time-consuming due to their reliance on rendering-based refinement approaches. To circumvent t…

2024

KP-RED: Exploiting Semantic Keypoints for Joint 3D Shape Retrieval and Deformation

CVPR 2024poster

In this paper we present KP-RED a unified KeyPoint-driven REtrieval and Deformation framework that takes object scans as input and jointly retrieves and deforms the most geometrically similar CAD models from a pre-processed database to tightly match the target. Unlike existing dense matching based m…

2024

LaPose: Laplacian Mixture Shape Modeling for RGB-Based Category-Level Object Pose Estimation

ECCV 2024poster

"While RGBD-based methods for category-level object pose estimation hold promise, their reliance on depth data limits their applicability in diverse scenarios. In response, recent efforts have turned to RGB-based methods; however, they face significant challenges stemming from the absence of depth i…

2024

MOHO: Learning Single-view Hand-held Object Reconstruction with Multi-view Occlusion-Aware Supervision

CVPR 2024poster

Previous works concerning single-view hand-held object reconstruction typically rely on supervision from 3D ground-truth models which are hard to collect in real world. In contrast readily accessible hand-object videos offer a promising training data source but they only give heavily occluded object…

Cited by 11SourcePDFScholar
2024

SG-Bot: Object Rearrangement via Coarse-to-Fine Robotic Imagination on Scene Graphs

ICRA 2024poster

Object rearrangement is pivotal in robotic-environment interactions, representing a significant capability in embodied AI. In this paper, we present SG-Bot, a novel rearrangement framework that utilizes a coarse-to-fine scheme with a scene graph as the scene representation. Unlike previous methods t…

Cited by 25SourceScholar
2024

SecondPose: SE(3)-Consistent Dual-Stream Feature Fusion for Category-Level Pose Estimation

CVPR 2024poster

Category-level object pose estimation aiming to predict the 6D pose and 3D size of objects from known categories typically struggles with large intra-class shape variation. Existing works utilizing mean shapes often fall short of capturing this variation. To address this issue we present SecondPose…

2024

ShapeMatcher: Self-Supervised Joint Shape Canonicalization Segmentation Retrieval and Deformation

CVPR 2024poster

In this paper we present ShapeMatcher a unified self-supervised learning framework for joint shape canonicalization segmentation retrieval and deformation. Given a partially-observed object in an arbitrary pose we first canonicalize the object by extracting point-wise affine invariant features disen…

2023

CommonScenes: Generating Commonsense 3D Indoor Scenes with Scene Graph Diffusion

NeurIPS 2023poster

Controllable scene synthesis aims to create interactive environments for numerous industrial use cases. Scene graphs provide a highly suitable interface to facilitate these applications by abstracting the scene context in a compact manner. Existing methods, reliant on retrieval from extensive databa…

2023

DDF-HO: Hand-Held Object Reconstruction via Conditional Directed Distance Field

NeurIPS 2023poster

Reconstructing hand-held objects from a single RGB image is an important and challenging problem. Existing works utilizing Signed Distance Fields (SDF) reveal limitations in comprehensively capturing the complex hand-object interactions, since SDF is only reliable within the proximity of the target…

2023

IPCC-TP: Utilizing Incremental Pearson Correlation Coefficient for Joint Multi-Agent Trajectory Prediction

CVPR 2023poster

Reliable multi-agent trajectory prediction is crucial for the safe planning and control of autonomous systems. Compared with single-agent cases, the major challenge in simultaneously processing multiple agents lies in modeling complex social interactions caused by various driving intentions and road…

Cited by 19SourcePDFScholar
2023

MonoGraspNet: 6-DoF Grasping with a Single RGB Image

ICRA 2023poster

6-DoF robotic grasping is a long-lasting but un-solved problem. Recent methods utilize strong 3D networks to extract geometric grasping representations from depth sensors, demonstrating superior accuracy on common objects but performing unsatisfactorily on photometrically challenging objects, e.g.,…

Cited by 39SourceScholar
2023

OPA-3D: Occlusion-Aware Pixel-Wise Aggregation for Monocular 3D Object Detection

RA-L 2023

Monocular 3D object detection has recently made a significant leap forward thanks to the use of pre-trained depth estimators for pseudo-LiDAR recovery. Yet, such two-stage methods typically suffer from overfitting and are incapable of explicitly encapsulating the geometric relation between depth and

Cited by 39SourceScholar
2023

Self-Supervised Category-Level 6D Object Pose Estimation With Optical Flow Consistency

RA-L 2023

Category-level 6D object pose estimation aims at determining the pose of an object of a given category. Most current state-of-the-art methods require a significant amount of real training data to supervise their models. Moreover, annotating the 6D pose is very time consuming, error-prone, and it doe

Cited by 19SourceScholar
2023

U-RED: Unsupervised 3D Shape Retrieval and Deformation for Partial Point Clouds

ICCV 2023poster

In this paper, we propose U-RED, an Unsupervised shape REtrieval and Deformation pipeline that takes an arbitrary object observation as input, typically captured by RGB images or scans, and jointly retrieves and deforms the geometrically similar CAD models from a pre-established database to tightly…

Cited by 3PDFcodeScholar
2022

GPV-Pose: Category-Level Object Pose Estimation via Geometry-Guided Point-Wise Voting

CVPR 2022poster

While 6D object pose estimation has recently made a huge leap forward, most methods can still only handle a single or a handful of different objects, which limits their applications. To circumvent this problem, category-level object pose estimation has recently been revamped, which aims at predictin…

Cited by 147PDFcodeScholar
2022

RBP-Pose: Residual Bounding Box Projection for Category-Level Pose Estimation

ECCV 2022poster

"Category-level object pose estimation aims to predict the 6D pose as well as the 3D metric size of previously unseen objects from a known set of categories. Recent methods harness shape prior adaptation to map the observed point cloud into the canonical space and apply Umeyama’s algorithm to recove…

2022

SSP-Pose: Symmetry-Aware Shape Prior Deformation for Direct Category-Level Object Pose Estimation

IROS 2022poster

Category-level pose estimation is a challenging problem due to intra-class shape variations. Recent methods deform pre-computed shape priors to map the observed point cloud into the normalized object coordinate space and then retrieve the pose via post-processing, i.e., Umeyama's Algorithm. The shor…

Cited by 41SourceScholar
2021

SO-Pose: Exploiting Self-Occlusion for Direct 6D Pose Estimation

ICCV 2021poster

Directly regressing all 6 degrees-of-freedom (6DoF) for the object pose (i.e. the 3D rotation and translation) in a cluttered environment from a single RGB image is a challenging problem. While end-to-end methods have recently demonstrated promising results at high efficiency, they are still inferio…

Cited by 160PDFcodeScholar
2020

A Unified Framework for Piecewise Semantic Reconstruction in Dynamic Scenes via Exploiting Superpixel Relations

ICRA 2020poster

This paper presents a novel framework for dense piecewise semantic reconstruction in dynamic scenes containing complex background and moving objects via exploiting superpixel relations. We utilize two kinds of superpixel relations: motion relations and spatial relations, each having three subcategor…

Cited by 10SourceScholar
2019

Monocular Piecewise Depth Estimation in Dynamic Scenes by Exploiting Superpixel Relations

ICCV 2019poster

In this paper, we propose a novel and specially designed method for piecewise dense monocular depth estimation in dynamic scenes. We utilize spatial relations between neighboring superpixels to solve the inherent relative scale ambiguity (RSA) problem and smooth the depth map. However, directly esti…

Cited by 8PDFScholar