← Search

Zizhao Wang

16 accepted papers

2026

Factored Latent Action World Models

ICML 2026poster

Learning latent actions from action-free video has emerged as a powerful paradigm for scaling up controllable world model learning. Latent actions provide a natural interface for users to iteratively generate and manipulate videos. However, most existing approaches rely on monolithic inverse and for…

Cited by 0SourceScholar
2025

Dyn-O: Building Structured World Models with Object-Centric Representations

NeurIPS 2025poster

World models aim to capture the dynamics of the environment, enabling agents to predict and plan for future states. In most scenarios of interest, the dynamics are highly centered on interactions among objects within the environment. This motivates the development of world models that operate on obj…

Cited by 0SourcecodeScholar
2025

Dyna-LfLH: Learning Agile Navigation in Dynamic Environments from Learned Hallucination

IROS 2025

This paper introduces Dynamic Learning from Learned Hallucination (Dyna-LfLH), a self-supervised method for training motion planners to navigate environments with dense and dynamic obstacles. Classical planners struggle with dense, unpredictable obstacles due to limited computation, while learning-b

Cited by 4SourceScholar
2024

Building Minimal and Reusable Causal State Abstractions for Reinforcement Learning

AAAI 2024technical

Two desiderata of reinforcement learning (RL) algorithms are the ability to learn from relatively little experience and the ability to learn policies that generalize to a range of problem specifications. In factored state spaces, one approach towards achieving both goals is to learn state abstracti…

Cited by 9SourcePDFScholar
2024

Disentangled Unsupervised Skill Discovery for Efficient Hierarchical Reinforcement Learning

NeurIPS 2024poster

A hallmark of intelligent agents is the ability to learn reusable skills purely from unsupervised interaction with the environment. However, existing unsupervised skill discovery methods often learn entangled skills where one skill variable simultaneously influences many entities in the environment,…

Cited by 2SourcePDFScholar
2024

SkiLD: Unsupervised Skill Discovery Guided by Factor Interactions

NeurIPS 2024poster

Unsupervised skill discovery carries the promise that an intelligent agent can learn reusable skills through autonomous, reward-free interactions with environments. Existing unsupervised skill discovery methods learn skills by encouraging distinguishable behaviors that cover diverse states. However,…

Cited by 1SourcePDFScholar
2022

Causal Dynamics Learning for Task-Independent State Abstraction

ICML 2022oral

Learning dynamics models accurately is an important goal for Model-Based Reinforcement Learning (MBRL), but most MBRL methods learn a dense dynamics model which is vulnerable to spurious correlations and therefore generalizes poorly to unseen states. In this paper, we introduce Causal Dynamics Learn…

2022

Learning to Correct Mistakes: Backjumping in Long-Horizon Task and Motion Planning

CoRL 2022poster

As robots become increasingly capable of manipulation and long-term autonomy, long-horizon task and motion planning problems are becoming increasingly important. A key challenge in such problems is that early actions in the plan may make future actions infeasible. When reaching a dead-end in the se…

Cited by 6SourceScholar
2021

APPLI: Adaptive Planner Parameter Learning From Interventions

ICRA 2021poster

While classical autonomous navigation systems can typically move robots from one point to another safely and in a collision-free manner, these systems may fail or produce suboptimal behavior in certain scenarios. The current practice in such scenarios is to manually re-tune the system’s parameters,…

Cited by 57SourceScholar
2021

APPLR: Adaptive Planner Parameter Learning from Reinforcement

ICRA 2021poster

Classical navigation systems typically operate using a fixed set of hand-picked parameters (e.g. maximum speed, sampling rate, inflation radius, etc.) and require heavy expert re-tuning in order to work in new environments. To mitigate this requirement, it has been proposed to learn parameters for d…

Cited by 61SourceScholar
2021

From Agile Ground to Aerial Navigation: Learning from Learned Hallucination

IROS 2021poster

This paper presents a self-supervised Learning from Learned Hallucination (LfLH) method to learn fast and reactive motion planners for ground and aerial robots to navigate through highly constrained environments. The recent Learning from Hallucination (LfH) paradigm for autonomous navigation execute…

Cited by 40SourceScholar
2020

Accelerated Robot Learning via Human Brain Signals

ICRA 2020poster

In reinforcement learning (RL), sparse rewards are a natural way to specify the task to be learned. However, most RL algorithms struggle to learn in this setting since the learning signal is mostly zeros. In contrast, humans are good at assessing and predicting the future consequences of actions and…

Cited by 25SourceScholar