← Search

Shenghua Wan

10 accepted papers

2026

Learning to Be Uncertain: Pre-training World Models with Horizon-Calibrated Uncertainty

ICLR 2026poster

Pre-training world models on large, action-free video datasets offers a promising path toward generalist agents, but a fundamental flaw undermines this paradigm. Prevailing methods train models to predict a single, deterministic future, an objective that is ill-posed for inherently stochastic enviro…

Cited by 0SourceScholar
2026

Multi-view Consistent Latent Action Learning for World Modeling and Control

ICML 2026poster

The scalability of world models is currently bottlenecked by the scarcity of action annotations. While self-supervised latent action learning offers a potential solution, existing single-view paradigms—relying on information bottlenecks or Vector Quantization (VQ)—often conflate superficial 2D pixel…

Cited by 0SourceScholar
2025

FOUNDER: Grounding Foundation Models in World Models for Open-Ended Embodied Decision Making

ICML 2025poster

Foundation Models (FMs) and World Models (WMs) offer complementary strengths in task generalization at different levels. In this work, we propose FOUNDER, a framework that integrates the generalizable knowledge embedded in FMs with the dynamic modeling capabilities of WMs to enable open-ended task s…

Cited by 0SourcePDFScholar
2025

Leveraging Conditional Dependence for Efficient World Model Denoising

NeurIPS 2025poster

Effective denoising is critical for managing complex visual inputs contaminated with noisy distractors in model-based reinforcement learning (RL). Current methods often oversimplify the decomposition of observations by neglecting the conditional dependence between task-relevant and task-irrelevant c…

Cited by 0SourceScholar
2025

Reward Models in Deep Reinforcement Learning: A Survey

IJCAI 2025

In reinforcement learning (RL), agents continually interact with the environment and use the feedback to refine their behavior. To guide policy optimization, reward models are introduced as proxies of the desired objectives, such that when the agent maximizes the accumulated reward, it also fulfills

Cited by 0SourcePDFScholar
2024

AD3: Implicit Action is the Key for World Models to Distinguish the Diverse Visual Distractors

ICML 2024poster

Model-based methods have significantly contributed to distinguishing task-irrelevant distractors for visual control. However, prior research has primarily focused on heterogeneous distractors like noisy background videos, leaving homogeneous distractors that closely resemble controllable agents larg…

Cited by 3SourcePDFScholar
2024

Leveraging Separated World Model for Exploration in Visually Distracted Environments

NeurIPS 2024poster

Model-based unsupervised reinforcement learning (URL) has gained prominence for reducing environment interactions and learning general skills using intrinsic rewards. However, distractors in observations can severely affect intrinsic reward estimation, leading to a biased exploration process, especi…

Cited by 1SourcePDFScholar
2024

MOSER: Learning Sensory Policy for Task-specific Viewpoint via View-conditional World Model

IJCAI 2024poster

Reinforcement learning from visual observations is a challenging problem with many real-world applications. Existing algorithms mostly rely on a single observation from a well-designed fixed camera that requires human knowledge. Recent studies learn from different viewpoints with multiple fixed came…

Cited by 0SourcePDFScholar
2024

SeMOPO: Learning High-quality Model and Policy from Low-quality Offline Visual Datasets

ICML 2024poster

Model-based offline reinforcement Learning (RL) is a promising approach that leverages existing data effectively in many real-world applications, especially those involving high-dimensional inputs like images and videos. To alleviate the distribution shift issue in offline RL, existing model-based m…

Cited by 0SourcePDFScholar
2023

SeMAIL: Eliminating Distractors in Visual Imitation via Separated Models

ICML 2023poster

Model-based imitation learning (MBIL) is a popular reinforcement learning method that improves sample efficiency on high-dimension input sources, such as images and videos. Following the convention of MBIL research, existing algorithms are highly deceptive by task-irrelevant information, especially…

Cited by 7SourcePDFScholar