← Search

Weidong Huang

6 accepted papers

2026

Towards Bridging the Gap between Large-Scale Pretraining and Efficient Finetuning for Humanoid Control

ICLR 2026poster

Reinforcement learning (RL) is widely used for humanoid control, with on-policy methods such as Proximal Policy Optimization (PPO) enabling robust training via large-scale parallel simulation and, in some cases, zero-shot deployment to real robots. However, the low sample efficiency of on-policy alg…

Cited by 0SourcecodeScholar
2025

Differentiable Information Enhanced Model-Based Reinforcement Learning

AAAI 2025technical

Differentiable environments have heralded new possibilities for learning control policies by offering rich differentiable information that facilitates gradient-based methods. In comparison to prevailing model-free reinforcement learning approaches, model-based reinforcement learning (MBRL) methods e…

Cited by 0SourcePDFScholar
2025

World Models Should Prioritize the Unification of Physical and Social Dynamics

NeurIPS 2025poster

World models, which explicitly learn environmental dynamics to lay the foundation for planning, reasoning, and decision-making, are rapidly advancing in predicting both physical dynamics and aspects of social behavior, yet predominantly in separate silos. This division results in a systemic failure…

Cited by 0SourceScholar
2024

SafeDreamer: Safe Reinforcement Learning with World Models

ICLR 2024poster

The deployment of Reinforcement Learning (RL) in real-world applications is constrained by its failure to satisfy safety criteria. Existing Safe Reinforcement Learning (SafeRL) methods, which rely on cost functions to enforce safety, often fail to achieve zero-cost performance in complex scenarios,…

2023

Safety Gymnasium: A Unified Safe Reinforcement Learning Benchmark

NeurIPS 2023poster

Artificial intelligence (AI) systems possess significant potential to drive societal progress. However, their deployment often faces obstacles due to substantial safety concerns. Safe reinforcement learning (SafeRL) emerges as a solution to optimize policies while simultaneously adhering to multiple…

Cited by 73SourcePDFScholar
2018

Error Compensated Quantized SGD and its Applications to Large-scale Distributed Optimization

ICML 2018oral

Large-scale distributed optimization is of great importance in various applications. For data-parallel based distributed learning, the inter-node gradient communication often becomes the performance bottleneck. In this paper, we propose the error compensated quantized stochastic gradient descent alg…

Cited by 286SourcePDFScholar