← Search

Jingliang Duan

12 accepted papers

2026

Mind Your Entropy: From Maximum Entropy to Trajectory Entropy-Constrained RL

ICML 2026poster

Maximum entropy has become a mainstream off-policy reinforcement learning (RL) framework for balancing exploitation and exploration. However, two bottlenecks still limit further performance gains: (1) non-stationary Q-value estimation stemming from the joint injection of entropy and the concurrent u…

Cited by 0SourceScholar
2026

Taming the Aleatoric Impulse in Off-Policy Reinforcement Learning

ICML 2026poster

Off-policy reinforcement learning is vulnerable to overestimation bias, which is rooted in the total value uncertainty. However, existing methods typically misaddress this by targeting the epistemic component, neglecting the aleatoric component. We identify for the first time that this oversight fai…

Cited by 0SourceScholar
2025

LipsNet++: Unifying Filter and Controller into a Policy Network

ICML 2025spotlight

Deep reinforcement learning (RL) is effective for decision-making and control tasks like autonomous driving and embodied AI. However, RL policies often suffer from the action fluctuation problem in real-world applications, resulting in severe actuator wear, safety risk, and performance degradation.…

2025

ODE-based Smoothing Neural Network for Reinforcement Learning Tasks

ICLR 2025spotlight

The smoothness of control actions is a significant challenge faced by deep reinforcement learning (RL) techniques in solving optimal control problems. Existing RL-trained policies tend to produce non-smooth actions due to high-frequency input noise and unconstrained Lipschitz constants in neural net…

Cited by 0SourcePDFScholar
2025

Off-policy Reinforcement Learning with Model-based Exploration Augmentation

NeurIPS 2025poster

Exploration is crucial in Reinforcement Learning (RL) as it enables the agent to understand the environment for better decision-making. Existing exploration methods fall into two paradigms: active exploration, which injects stochasticity into the policy but struggles in high-dimensional environments…

Cited by 0SourceScholar
2025

Physics Informed Neural Pose Estimation for Real-Time Shape Reconstruction of Soft Continuum Robots

RA-L 2025

Soft continuum robots are increasingly valued for their remarkable flexibility, but accurate shape reconstruction remains challenging due to their infinite degrees of freedom and high nonlinearity. Existing approaches often rely on either simplified-curvature statics equations for physical derivatio

Cited by 4SourceScholar
2024

Diffusion Actor-Critic with Entropy Regulator

NeurIPS 2024poster

Reinforcement learning (RL) has proven highly effective in addressing complex decision-making and control tasks. However, in most traditional RL algorithms, the policy is typically parameterized as a diagonal Gaussian distribution with learned mean and variance, which constrains their capability to…

2024

Global Optimality of Single-Timescale Actor-Critic under Continuous State-Action Space: A Study on Linear Quadratic Regulator

IJCAI 2024poster

Actor-critic methods have achieved state-of-the-art performance in various challenging tasks. However, theoretical understandings of their performance remain elusive and challenging. Existing studies mostly focus on practically uncommon variants such as double-loop or two-timescale stepsize actor-cr…

Cited by 1SourcePDFScholar
2024

Guiding Reinforcement Learning with Incomplete System Dynamics

IROS 2024poster

Model-free reinforcement learning (RL) is inherently a reactive method, operating under the assumption that it starts with no prior knowledge of the system and entirely depends on trial-and-error for learning. This approach faces several challenges, such as poor sample efficiency, generalization, an…

Cited by 2SourceScholar
2023

Global Convergence of Two-Timescale Actor-Critic for Solving Linear Quadratic Regulator

AAAI 2023technical

The actor-critic (AC) reinforcement learning algorithms have been the powerhouse behind many challenging applications. Nevertheless, its convergence is fragile in general. To study its instability, existing works mostly consider the uncommon double-loop variant or basic models with finite state and…

Cited by 12SourcePDFScholar
2023

LipsNet: A Smooth and Robust Neural Network with Adaptive Lipschitz Constant for High Accuracy Optimal Control

ICML 2023poster

Deep reinforcement learning (RL) is a powerful approach for solving optimal control problems. However, RL-trained policies often suffer from the action fluctuation problem, where the consecutive actions significantly differ despite only slight state variations. This problem results in mechanical com…

Cited by 17SourcePDFScholar