← Search

Philipp Wu

9 accepted papers

2026

SARM: Stage-Aware Reward Modeling for Long Horizon Robot Manipulation

ICLR 2026poster

Large-scale robot learning has made progress on complex manipulation tasks, yet long-horizon, contact-rich problems—especially those involving deformable objects—remain challenging due to inconsistent demonstration quality. We propose a stage-aware, video-based reward modeling framework that jointly…

Cited by 0SourcecodeScholar
2025

The Sound of Simulation: Learning Multimodal Sim-to-Real Robot Policies with Generative Audio

CoRL 2025oral

Robots must integrate multiple sensory modalities to act effectively in the real world. Yet, learning such multimodal policies at scale remains challenging. Simulation offers a viable solution, but while vision has benefited from high-fidelity simulators, other modalities (e.g. sound) can be notorio…

Cited by 0SourceScholar
2024

From LLMs to Actions: Latent Codes as Bridges in Hierarchical Robot Control

IROS 2024poster

Hierarchical control for robotics has long been plagued by the need to have a well defined interface layer to communicate between high-level task planners and low-level policies. With the advent of LLMs, language has been emerging as a prospective interface layer. However, this has several limitatio…

Cited by 11SourceScholar
2024

GELLO: A General, Low-Cost, and Intuitive Teleoperation Framework for Robot Manipulators

IROS 2024poster

Humans can teleoperate robots to accomplish complex manipulation tasks. Imitation learning has emerged as a powerful framework that leverages human teleoperated demonstrations to teach robots new skills. However, the performance of the learned policies is bottlenecked by the quality, scale, and vari…

Cited by 106SourcecodeScholar
2023

Masked Trajectory Models for Prediction, Representation, and Control

ICML 2023poster

We introduce Masked Trajectory Models (MTM) as a generic abstraction for sequential decision making. MTM takes a trajectory, such as a state-action sequence, and aims to reconstruct the trajectory conditioned on random subsets of the same trajectory. By training with a highly randomized masking patt…

2023

RoboPianist: Dexterous Piano Playing with Deep Reinforcement Learning

CoRL 2023poster

Replicating human-like dexterity in robot hands represents one of the largest open problems in robotics. Reinforcement learning is a promising approach that has achieved impressive progress in the last few years; however, the class of problems it has typically addressed corresponds to a rather narro…

Cited by 47SourcecodeScholar
2022

DayDreamer: World Models for Physical Robot Learning

CoRL 2022poster

To solve tasks in complex environments, robots need to learn from experience. Deep reinforcement learning is a common approach to robot learning but requires a large amount of trial and error to learn, limiting its deployment in the physical world. As a consequence, many advances in robot learning r…

Cited by 328SourcecodeScholar
2021

Replay Overshooting: Learning Stochastic Latent Dynamics with the Extended Kalman Filter

ICRA 2021poster

This paper presents replay overshooting (RO), an algorithm that uses properties of the extended Kalman filter (EKF) to learn nonlinear stochastic latent dynamics models suitable for long-horizon prediction. We build upon overshooting methods used to train other prediction models and recover a novel…

Cited by 11SourceScholar
2019

Quasi-Direct Drive for Low-Cost Compliant Robotic Manipulation

ICRA 2019poster

Robots must cost less and be force-controlled to enable widespread, safe deployment in unconstrained human environments. We propose Quasi-Direct Drive actuation as a capable paradigm for robotic force-controlled manipulation in human environments at low-cost. Our prototype - Blue - is a human scale…

Cited by 115SourcecodeScholar