← Search

Alessandro Montenegro

6 accepted papers

2026

Mind Your Steps: A General Learning Framework for Accurate Humanoid Foothold Tracking

RSS 2026poster

Enabling humanoid robots to operate in complex, dynamic environments remains a critical challenge, fundamentally limited by the ability to navigate robustly, safely, and accurately. While reinforcement learning with velocity-commanded policies has achieved remarkable robustness in humanoid locomotio…

Cited by 0SourceScholar
2026

Reusing Trajectories in Policy Gradients Enables Fast Convergence

ICML 2026poster

*Policy gradient* (PG) methods are a class of effective *reinforcement learning* algorithms, particularly when dealing with continuous control problems. They rely on fresh *on-policy* data, making them sample-inefficient and requiring $\mathcal{O}(\epsilon^{-2})$ trajectories to reach an $\epsilon$-…

Cited by 0SourceScholar
2025

Convergence Analysis of Policy Gradient Methods with Dynamic Stochasticity

ICML 2025poster

*Policy gradient* (PG) methods are effective *reinforcement learning* (RL) approaches, particularly for continuous problems. While they optimize stochastic (hyper)policies via action- or parameter-space exploration, real-world applications often require deterministic policies. Existing PG convergenc…

Cited by 0SourcePDFScholar
2024

Best Arm Identification for Stochastic Rising Bandits

ICML 2024spotlight

Stochastic Rising Bandits (SRBs) model sequential decision-making problems in which the expected reward of the available options increases every time they are selected. This setting captures a wide range of scenarios in which the available options are learning entities whose performance improves (in…

2024

Last-Iterate Global Convergence of Policy Gradients for Constrained Reinforcement Learning

NeurIPS 2024poster

*Constrained Reinforcement Learning* (CRL) tackles sequential decision-making problems where agents are required to achieve goals by maximizing the expected return while meeting domain-specific constraints, which are often formulated on expected costs. In this setting, *policy-based* methods are wid…

Cited by 2SourcePDFScholar
2024

Learning Optimal Deterministic Policies with Stochastic Policy Gradients

ICML 2024spotlight

Policy gradient (PG) methods are successful approaches to deal with continuous reinforcement learning (RL) problems. They learn stochastic parametric (hyper)policies by either exploring in the space of actions or in the space of parameters. Stochastic controllers, however, are often undesirable from…

Cited by 2SourcePDFScholar