← Search

Jianglong Ye

13 accepted papers

2026

Cross-Hand Latent Representation for Vision-Language-Action Models

CVPR 2026

Dexterous manipulation is essential for real-world robot autonomy, mirroring the central role of human hand coordination in daily activity. Humans rely on rich multimodal perception--vision, sound, and language-guided intent--to perform dexterous actions, motivating vision-based, language-conditione

Cited by 0SourceScholar
2025

Co-Design of Soft Gripper with Neural Physics

CoRL 2025poster

For robot manipulation, both the controller and end-effector design are crucial. Compared with rigid grippers, soft grippers are more generalizable by deforming to different geometries, but designing such a gripper and finding its grasp pose remains challenging. In this paper, we propose a co-design…

Cited by 0SourceScholar
2025

Dex1B: Learning with 1B Demonstrations for Dexterous Manipulation

RSS 2025poster

Generating large-scale demonstrations for dexterous manipulation remains a challenging problem, and various approaches have been proposed in recent years to address it. Among these, generative models have emerged as a promising paradigm, enabling the efficient generation of diverse and plausible dem…

Cited by 0PDFScholar
2025

Learning Generalizable Feature Fields for Mobile Manipulation

IROS 2025

An open problem in mobile manipulation is how to represent objects and scenes in a unified manner so that robots can use both for navigation and manipulation. The latter requires capturing intricate geometry while understanding fine-grained semantics, whereas the former involves capturing the comple

Cited by 49SourceScholar
2024

MVDream: Multi-view Diffusion for 3D Generation

ICLR 2024poster

We introduce MVDream, a diffusion model that is able to generate consistent multi-view images from a given text prompt. Learning from both 2D and 3D data, a multi-view diffusion model can achieve the generalizability of 2D diffusion models and the consistency of 3D renderings. We demonstrate that su…

Cited by 630SourcePDFScholar
2023

FeatureNeRF: Learning Generalizable NeRFs by Distilling Foundation Models

ICCV 2023poster

Recent works on generalizable NeRFs have shown promising results on novel view synthesis from single or few images. However, such models have rarely been applied on other downstream tasks beyond synthesis such as semantic understanding and parsing. In this paper, we propose a novel framework named F…

Cited by 46PDFcodeScholar
2023

GNFactor: Multi-Task Real Robot Learning with Generalizable Neural Feature Fields

CoRL 2023oral

It is a long-standing problem in robotics to develop agents capable of executing diverse manipulation tasks from visual observations in unstructured real-world environments. To achieve this goal, the robot will need to have a comprehensive understanding of the 3D structure and semantics of the scen…

Cited by 88SourcecodeScholar
2023

Learning Continuous Grasping Function With a Dexterous Hand From Human Demonstrations

RA-L 2023

We propose to learn to generate grasping motion for manipulation with a dexterous hand using implicit functions. With continuous time inputs, the model can generate a continuous and smooth grasping plan. We name the proposed model Continuous Grasping Function (CGF). CGF is learned via generative mod

Cited by 75SourcecodeScholar
2022

GIFS: Neural Implicit Function for General Shape Representation

CVPR 2022poster

Recent development of neural implicit function has shown tremendous success on high-quality 3D shape reconstruction. However, most works divide the space into inside and outside of the shape, which limits their representing power to single-layer and watertight shapes. This limitation leads to tediou…

Cited by 75PDFcodeScholar
2022

Online Adaptation for Implicit Object Tracking and Shape Reconstruction in the Wild

RA-L 2022

Tracking and reconstructing 3D objects from cluttered scenes are the key components for computer vision, robotics and autonomous driving systems. While recent progress in implicit function has shown encouraging results on high-quality 3D shape reconstruction, it is still very challenging to generali

Cited by 9SourcecodeScholar