← Search

Harry Mead

2 accepted papers

2025

Improving Regret Approximation for Unsupervised Dynamic Environment Generation

NeurIPS 2025poster

Unsupervised Environment Design (UED) seeks to automatically generate training curricula for reinforcement learning (RL) agents, with the goal of improving generalisation and zero-shot performance. However, designing effective curricula remains a difficult problem, particularly in settings where sma…

Cited by 0SourcecodeScholar
2025

Return Capping: Sample Efficient CVaR Policy Gradient Optimisation

ICML 2025poster

When optimising for conditional value at risk (CVaR) using policy gradients (PG), current methods rely on discarding a large proportion of trajectories, resulting in poor sample efficiency. We propose a reformulation of the CVaR optimisation problem by capping the total return of trajectories used…