← Search

Haoru Xue

10 accepted papers

2026

CaP-X: A Framework for Benchmarking and Improving Coding Agents for Robot Manipulation

ICML 2026poster

“Code-as-Policy” considers how executable code can complement data-intensive Vision-LanguageAction (VLA) methods, yet their effectiveness as autonomous controllers for embodied manipulation remains underexplored. We present CaPX, an open-access framework for systematically studying Code-as-Policy ag…

Cited by 0SourcecodeScholar
2026

Learning to Grasp Anything By Playing with Random Toys

ICLR 2026poster

Robotic manipulation policies often struggle to generalize to novel objects, limiting their real-world utility. In contrast, cognitive science suggests that children develop generalizable dexterous manipulation skills by mastering a small set of simple toys and then applying that knowledge to more c…

Cited by 0SourceScholar
2026

Opening the Sim-to-Real Door for Humanoid Pixel-to-Action Policy Transfer

CVPR 2026

Recent progress in GPU-accelerated, photorealistic simulation has opened a scalable data-generation path for robot learning, where massive physics and visual randomization allow policies to generalize beyond curated environments. Building on these advances, we develop a teacher-student-bootstrap lea

Cited by 0SourcecodeScholar
2026

Self-Improving Vision-Language-Action Models with Data Generation via Residual RL

ICLR 2026poster

Supervised fine-tuning (SFT) has become the de facto post-training strategy for large vision-language-action (VLA) models, but its reliance on costly human demonstrations limits scalability and generalization. We propose Probe, Learn, Distill (PLD), a plug-and-play framework that improves VLAs throu…

Cited by 0SourceScholar
2026

VIRAL: Visual Sim-to-Real at Scale for Humanoid Loco-Manipulation

CVPR 2026

A key barrier to the real-world deployment of humanoid robots is the lack of autonomous loco-manipulation skills. We introduce VIRAL, a visual sim-to-real framework that learns humanoid loco-manipulation entirely in simulation and deploys it zero-shot to real hardware. VIRAL follows a teacher-studen

Cited by 0SourcecodeScholar
2025

Agile Mobility with Rapid Online Adaptation via Meta-Learning and Uncertainty-Aware MPPI

ICRA 2025

Modern non-linear model-based controllers require an accurate physics model and model parameters to be able to control mobile robots at their limits. Also, due to surface slipping at high speeds, the friction parameters may continually change (like tire degradation in autonomous racing), and the con

Cited by 3SourceScholar
2025

AnyCar to Anywhere: Learning Universal Dynamics Model for Agile and Adaptive Mobility

ICRA 2025

Recent works in the robot learning community have successfully introduced generalist models capable of controlling various robot embodiments across a wide range of tasks, such as navigation and locomotion. However, achieving agile control, which pushes the limits of robotic performance, still relies

Cited by 25SourceScholar
2025

Full-Order Sampling-Based MPC for Torque-Level Locomotion Control via Diffusion-Style Annealing

ICRA 2025

Due to high dimensionality and non-convexity, real-time optimal control using full-order dynamics models for legged robots is challenging. Therefore, Nonlinear Model Predictive Control (NMPC) approaches are often limited to reduced-order models or local approximations. Sampling-based MPC has shown p

Cited by 69SourceScholar
2025

Pre-training Auto-regressive Robotic Models with 4D Representations

ICML 2025poster

Foundation models pre-trained on massive unlabeled datasets have revolutionized natural language and computer vision, exhibiting remarkable generalization capabilities, thus highlighting the importance of pre-training. Yet, efforts in robotics have struggled to achieve similar success, limited by ei…

Cited by 0SourcePDFScholar
2024

Learning Model Predictive Control with Error Dynamics Regression for Autonomous Racing

ICRA 2024poster

This work presents a novel Learning Model Predictive Control (LMPC) strategy for autonomous racing at the handling limit that can iteratively explore and learn unknown dynamics in high-speed operational domains. We start from existing LMPC formulations and modify the system dynamics learning method.…

Cited by 9SourcecodeScholar