← Search

Ziyang Tang

7 accepted papers

2025

Think on Your Feet: Seamless Transition Between Human-Like Locomotion in Response to Changing Commands

ICRA 2025

While it is relatively easier to train humanoid robots to mimic specific locomotion skills, it is more challenging to learn from various motions and adhere to continuously changing commands. These robots must accurately track motion instructions, seamlessly transition between a variety of movements,

Cited by 2SourceScholar
2024

Adapting Humanoid Locomotion over Challenging Terrain via Two-Phase Training

CoRL 2024poster

Humanoid robots are a key focus in robotics, with their capacity to navigate tough terrains being essential for many uses. While strides have been made, creating adaptable locomotion for complex environments is still tough. Recent progress in learning-based systems offers hope for robust legged loco…

Cited by 4SourceScholar
2021

Non-asymptotic Confidence Intervals of Off-policy Evaluation: Primal and Dual Bounds

ICLR 2021poster

Off-policy evaluation (OPE) is the task of estimating the expected reward of a given policy based on offline data previously collected under different policies. Therefore, OPE is a key step in applying reinforcement learning to real-world domains such as medical treatment, where interactive data col…

Cited by 16SourcePDFScholar
2020

Accountable Off-Policy Evaluation With Kernel Bellman Statistics

ICML 2020poster

We consider off-policy evaluation (OPE), which evaluates the performance of a new policy from observed data collected from previous experiments, without requiring the execution of the new policy. This finds important applications in areas with high execution cost or safety concerns, such as medical…

Cited by 49SourcePDFScholar
2020

Off-Policy Interval Estimation with Lipschitz Value Iteration

NeurIPS 2020poster

Off-policy evaluation provides an essential tool for evaluating the effects of different policies or treatments using only observed data. When applied to high-stakes scenarios such as medical diagnosis or financial decision-making, it is essential to provide provably correct upper and lower bounds o…

Cited by 5SourcePDFScholar
2019

Stein Variational Gradient Descent With Matrix-Valued Kernels

NeurIPS 2019poster

Stein variational gradient descent (SVGD) is a particle-based inference algorithm that leverages gradient information for efficient approximate inference. In this work, we enhance SVGD by leveraging preconditioning matrices, such as the Hessian and Fisher information matrix, to incorporate geometri…

2018

Breaking the Curse of Horizon: Infinite-Horizon Off-Policy Estimation

NeurIPS 2018spotlight

We consider the off-policy estimation problem of estimating the expected reward of a target policy using samples collected by a different behavior policy. Importance sampling (IS) has been a key technique to derive (nearly) unbiased estimators, but is known to suffer from an excessively high varianc…

Cited by 429SourcePDFScholar