← Search

Huaze Tang

8 accepted papers

2026

Policy Newton Algorithm in Reproducing Kernel Hilbert Space

ICLR 2026poster

Reinforcement learning (RL) policies represented in Reproducing Kernel Hilbert Spaces (RKHS) offer powerful representational capabilities. While second-order optimization methods like Newton's method demonstrate faster convergence than first-order approaches, current RKHS-based policy optimization r…

Cited by 0SourceScholar
2025

Distributional Decision Transformer: Risk-Sensitive Offline RL via Quantile-Based Critics and Stochastic Return

IROS 2025

Offline reinforcement learning faces a critical challenge in synthesizing high-reward trajectories from suboptimal datasets while robustly handling the stochasticity inherent in real-world decision-making. While combination of return-conditioned sequence models, such as Decision Transformers (DT), a

Cited by 0SourceScholar
2025

Mean-Field Aided QMIX: A Scalable and Flexible Q-Learning Approach for Large-Scale Agent Groups

ICASSP 2025accepted

Value decomposition methods are effective for multi-agent reinforcement learning (MARL), with QMIX being one of the most advanced. However, it struggles with scalability and flexibility in large-scale agent systems. The introduction of mean-field theory into MARL provides a potential solution to bot…

Cited by 0SourceScholar
2025

Reinforced Domain Selection for Continuous Domain Adaptation

ICASSP 2025accepted

Continuous Domain Adaptation (CDA) effectively bridges significant domain shifts by progressively adapting from the source domain through intermediate domains to the target domain. However, selecting intermediate domains without explicit metadata remains a substantial challenge that has not been ext…

Cited by 0SourceScholar
2025

Residual Kernel Policy Network: Enhancing Stability and Robustness in RKHS-Based Reinforcement Learning

ICLR 2025poster

Achieving optimal performance in reinforcement learning requires robust policies supported by training processes that ensure both sample efficiency and stability. Modeling the policy in reproducing kernel Hilbert space (RKHS) enables efficient exploration of local optimal solutions. However, the sta…

Cited by 0SourcePDFScholar
2024

M3ARL: Moment-Embedded Mean-Field Multi-Agent Reinforcement Learning for Continuous Action Space

ICASSP 2024accepted

Mean-field theory offers a promising solution to the scalability issues encountered in multi-agent reinforcement learning (MARL) within large-scale systems. However, most existing MARL algorithms based on mean-field theory are typically constrained to discrete action space. In continuous action spac…

Cited by 0SourceScholar
2023

Autonomous Swarm Robot Coordination via Mean-Field Control Embedding Multi-Agent Reinforcement Learning

IROS 2023poster

The learning approaches of designing a controller to guide the collective behavior of swarm robots have gained significant attention in recent years. However, the scalability of swarm robots and their inherent stochasticity complicate the control problem due to increasing complexity, unpredictabilit…

Cited by 4SourceScholar
2023

STEV: Stretchable Triboelectric E-skin enabled Proprioceptive Vibration Sensing for Soft Robot

ICRA 2023poster

Vibration perception is essential for robotic sensing and dynamic control. Nevertheless, due to the rigorous demand for sensor conformability and stretchability, enabling soft robots with proprioceptive vibration sensing remains challenging. This paper proposes a novel liquid metal-based stretchable…

Cited by 6SourceScholar