← Search

Sizhe Yang

10 accepted papers

2026

DexImit: Learning Bimanual Dexterous Manipulation from Monocular Human Videos

RSS 2026poster

Data scarcity fundamentally limits the generalization of bimanual dexterous manipulation, as real-world data collection for dexterous hands is expensive and labor-intensive. Human manipulation videos, as a direct carrier of manipulation knowledge, offer significant potential for scaling up robot lea…

Cited by 0SourceScholar
2026

MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence

ICLR 2026poster

Spatial intelligence is essential for multimodal large language models (MLLMs) operating in the complex physical world. Existing benchmarks, however, probe only single-image relations and thus fail to assess the multi-image spatial reasoning that real-world deployments demand. We introduce MMSI-Benc…

Cited by 0SourcecodeScholar
2026

One-Policy-Fits-All: Geometry-Aware Action Latents for Cross-Embodiment Manipulation

ICRA 2026poster

Cross-embodiment manipulation is crucial for enhancing the scalability of robot manipulation and reducing the high cost of data collection. However, the significant differences between embodiments, such as variations in action spaces and structural disparities, pose challenges for joint training acr…

2026

Robo3R: Enhancing Robotic Manipulation with Accurate Feed-Forward 3D Reconstruction

RSS 2026poster

3D spatial perception is fundamental to generalizable robotic manipulation, yet obtaining reliable, high-quality 3D geometry remains challenging. Depth sensors suffer from noise and material sensitivity, while existing reconstruction models lack the precision and metric consistency required for phys…

Cited by 0SourceScholar
2026

UltraDexGrasp: Learning Universal Dexterous Grasping for Bimanual Robots with Synthetic Data

ICRA 2026poster

Grasping is a fundamental capability for robots to interact with the physical world. Humans, equipped with two hands, autonomously select appropriate grasp strategies based on the shape, size, and weight of objects, enabling robust grasping and subsequent manipulation. In contrast, current robotic g…

2025

FCoDT-Net: A Novel Framework for High-Precision Medical Image Segmentation Using Contextual Distillation Transformer

ICASSP 2025accepted

Current methods in medical image semantic segmentation often rely on simple skip connections within U-shaped network structures. These approaches fail to bridge the semantic gap between the encoder and decoder module and do not fully exploit the rich contextual information among. The unused informat…

Cited by 0SourceScholar
2025

Novel Demonstration Generation with Gaussian Splatting Enables Robust One-Shot Manipulation

RSS 2025poster

Visuomotor policies learned through imitation learning methods often struggle to generalize to new visual domains due to the limited diversity of expert demonstrations, and collecting extensive real-world data is exhaustive. To address this challenge, we propose a novel demonstration generation app…

Cited by 1PDFScholar
2025

Predictive Inverse Dynamics Models are Scalable Learners for Robotic Manipulation

ICLR 2025oral

Current efforts to learn scalable policies in robotic manipulation primarily fall into two categories: one focuses on "action," which involves behavior cloning from extensive collections of robotic data, while the other emphasizes "vision," enhancing model generalization by pre-training representati…

2023

RL-ViGen: A Reinforcement Learning Benchmark for Visual Generalization

NeurIPS 2023poster

Visual Reinforcement Learning (Visual RL), coupled with high-dimensional observations, has consistently confronted the long-standing challenge of out-of-distribution generalization. Despite the focus on algorithms aimed at resolving visual generalization problems, we argue that the devil is in the e…