← Search

Xin Fei

5 accepted papers

2026

T(R, O) Grasp: Efficient Graph Diffusion of Robot-Object Spatial Transformation for Cross-Embodiment Dexterous Grasping

ICRA 2026poster

Dexterous grasping remains a central challenge in robotics due to the complexity of its high-dimensional state and action space. We introduce T(R,O) Grasp, a diffusion-based framework that efficiently generates accurate and diverse grasps across multiple robotic hands. At its core is the T(R,O) Grap…

2025

Scene Splatter: Momentum 3D Scene Generation from Single Image with Video Diffusion Model

CVPR 2025poster

In this paper, we propose Scene Splatter, a momentum-based paradigm for video diffusion to generate generic scenes from single image. Existing methods, which employ video generation models to synthesize novel views, suffer from limited video length and scene inconsistency, leading to artifacts and d…

Cited by 1SourcePDFScholar
2025

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models

NeurIPS 2025poster

Recent studies on Vision-Language-Action (VLA) models have shifted from the end-to-end action-generation paradigm toward a pipeline involving task planning followed by action generation, demonstrating improved performance on various complex, long-horizon manipulation tasks. However, existing approac…

Cited by 0SourceScholar
2024

Gaussian Graph Network: Learning Efficient and Generalizable Gaussian Representations from Multi-view Images

NeurIPS 2024poster

3D Gaussian Splatting (3DGS) has demonstrated impressive novel view synthesis performance. While conventional methods require per-scene optimization, more recently several feed-forward methods have been proposed to generate pixel-aligned Gaussian representations with a learnable network, which are g…

Cited by 1SourcePDFScholar
2024

GeoAuxNet: Towards Universal 3D Representation Learning for Multi-sensor Point Clouds

CVPR 2024poster

Point clouds captured by different sensors such as RGB-D cameras and LiDAR possess non-negligible domain gaps. Most existing methods design different network architectures and train separately on point clouds from various sensors. Typically point-based methods achieve outstanding performances on eve…