← Search

Haoyu Zhen

8 accepted papers

2026

GHOST: Hierarchical Sub-Goal Policies for Generalizing Robot Manipulation

RSS 2026poster

We present GHOST, a framework for learning visuomotor manipulation policies that generalize beyond the training distribution. GHOST factorizes control into (i) a high-level policy that predicts the next sub-goal as a distribution over 3D end-effector poses from multi-view RGB-D observations, and (ii…

Cited by 0SourceScholar
2025

Learning 4D Embodied World Models

ICCV 2025poster

This paper presents an effective approach for learning novel 4D embodied world models, which predict the dynamic evolution of 3D scenes over time in response to an embodied agent's actions, providing both spatial and temporal consistency. We propose to learn a 4D world model by training on RGB-DN (R…

2025

RapVerse: Coherent Vocals and Whole-Body Motion Generation from Text

ICCV 2025poster

In this work, we introduce a challenging task for simultaneously generating 3D holistic body motions and singing vocals directly from textual lyrics inputs, advancing beyond existing works that typically address these two modalities in isolation. To facilitate this, we first collect the RapVerse dat…

Cited by 0SourcePDFScholar
2024

3D-VLA: A 3D Vision-Language-Action Generative World Model

ICML 2024poster

Recent vision-language-action (VLA) models rely on 2D inputs, lacking integration with the broader realm of the 3D physical world. Furthermore, they perform action prediction by learning a direct mapping from perception to action, neglecting the vast dynamics of the world and the relations between a…

Cited by 77SourcePDFScholar
2023

3D-LLM: Injecting the 3D World into Large Language Models

NeurIPS 2023spotlight

Large language models (LLMs) and Vision-Language Models (VLMs) have been proved to excel at multiple tasks, such as commonsense reasoning. Powerful as these models can be, they are not grounded in the 3D physical world, which involves richer concepts such as spatial relationships, affordances, physi…

Cited by 328SourcePDFScholar
2023

CHORD: Category-level Hand-held Object Reconstruction via Shape Deformation

ICCV 2023poster

In daily life, humans utilize hands to manipulate objects. Modeling the shape of objects that are manipulated by the hand is essential for AI to comprehend daily tasks and to learn manipulation skills. However, previous approaches have encountered difficulties in reconstructing the precise shapes of…

Cited by 15PDFcodeScholar
2023

Relative Entropic Optimal Transport: a (Prior-aware) Matching Perspective to (Unbalanced) Classification

NeurIPS 2023poster

Classification is a fundamental problem in machine learning, and considerable efforts have been recently devoted to the demanding long-tailed setting due to its prevalence in nature. Departure from the Bayesian framework, this paper rethinks classification from a matching perspective by studying the…

2023

Understanding and Generalizing Contrastive Learning from the Inverse Optimal Transport Perspective

ICML 2023poster

Previous research on contrastive learning (CL) has primarily focused on pairwise views to learn representations by attracting positive samples and repelling negative ones. In this work, we aim to understand and generalize CL from a point set matching perspective, instead of the comparison between tw…

Cited by 19SourcePDFScholar