← Search

Minghui Zhu

16 accepted papers

2025

DSTE-Net: Dual-Scale Spatial-Temporal Excitation Network for Dynamic Gesture Recognition

RA-L 2025

Dynamic gesture recognition is a crucial technology for achieving natural human-computer interaction and holds broad application prospects in fields such as virtual reality, smart home, and digital entertainment. However, existing methods often lack a unified mechanism to simultaneously capture subt

Cited by 0SourceScholar
2025

Explainable Reinforcement Learning from Human Feedback to Improve Alignment

NeurIPS 2025poster

A common and effective strategy for humans to improve an unsatisfactory outcome in daily life is to find a cause of this outcome and correct the cause. In this paper, we investigate whether this human improvement strategy can be applied to improving reinforcement learning from human feedback (RLHF)…

Cited by 0SourceScholar
2025

Meta-Reinforcement Learning with Adaptation from Human Feedback via Preference-Order-Preserving Task Embedding

ICML 2025poster

This paper studies meta-reinforcement learning with adaptation from human feedback. It aims to pre-train a meta-model that can achieve few-shot adaptation for new tasks from human preference queries without relying on reward signals. To solve the problem, we propose the framework *adaptation via Pre…

Cited by 0SourcePDFScholar
2025

UTILITY: Utilizing Explainable Reinforcement Learning to Improve Reinforcement Learning

ICLR 2025poster

Reinforcement learning (RL) faces two challenges: (1) The RL agent lacks explainability. (2) The trained RL agent is, in many cases, non-optimal and even far from optimal. To address the first challenge, explainable reinforcement learning (XRL) is proposed to explain the decision-making of the RL ag…

Cited by 0SourcePDFScholar
2024

In-Trajectory Inverse Reinforcement Learning: Learn Incrementally Before an Ongoing Trajectory Terminates

NeurIPS 2024poster

Inverse reinforcement learning (IRL) aims to learn a reward function and a corresponding policy that best fit the demonstrated trajectories of an expert. However, current IRL works cannot learn incrementally from an ongoing trajectory because they have to wait to collect at least one complete trajec…

Cited by 3SourcePDFScholar
2024

Meta Inverse Constrained Reinforcement Learning: Convergence Guarantee and Generalization Analysis

ICLR 2024poster

This paper considers the problem of learning the reward function and constraints of an expert from few demonstrations. This problem can be considered as a meta-learning problem where we first learn meta-priors over reward functions and constraints from other distinct but related tasks and then adapt…

Cited by 22SourcePDFScholar
2024

Meta-Reinforcement Learning with Universal Policy Adaptation: Provable Near-Optimality under All-task Optimum Comparator

NeurIPS 2024poster

Meta-reinforcement learning (Meta-RL) has attracted attention due to its capability to enhance reinforcement learning (RL) algorithms, in terms of data efficiency and generalizability. In this paper, we develop a bilevel optimization framework for meta-RL (BO-MRL) to learn the meta-prior for task-sp…

Cited by 1SourcePDFScholar
2022

Byzantine-tolerant federated Gaussian process regression for streaming data

NeurIPS 2022accept

In this paper, we consider Byzantine-tolerant federated learning for streaming data using Gaussian process regression (GPR). In particular, a cloud and a group of agents aim to collaboratively learn a latent function where some agents are subject to Byzantine attacks. We develop a Byzantine-tolerant…

Cited by 5SourcePDFScholar
2020

Data-driven Distributed State Estimation and Behavior Modeling in Sensor Networks

IROS 2020poster

Nowadays, the prevalence of sensor networks has enabled tracking of the states of dynamic objects for a wide spectrum of applications from autonomous driving to environmental monitoring and urban planning. However, tracking realworld objects often faces two key challenges: First, due to the limitati…

Cited by 7SourceScholar