← Search

Esther Derman

11 accepted papers

2026

Discrete Compositional Generation via General Soft Operators and Robust Reinforcement Learning

ICLR 2026poster

A major bottleneck in scientific discovery consists of narrowing an exponentially large set of objects, such as proteins or molecules, to a small set of promising candidates with desirable properties. While this process can rely on expert knowledge, recent methods leverage reinforcement learning (RL…

Cited by 0SourceScholar
2026

Long-Horizon Model-Based Offline Reinforcement Learning Without Conservatism

ICML 2026poster

Popular offline reinforcement learning (RL) methods rely on conservatism, penalizing out-of-dataset actions or restricting rollout horizons. We question the universality of this principle and revisit a complementary Bayesian perspective. By modeling a posterior over plausible world models and traini…

Cited by 0SourceScholar
2026

Reward Redistribution for CVaR MDPs using a Bellman Operator on L-infinity

ICML 2026spotlight

Tail-end risk measures such as static conditional value-at-risk (CVaR) are used in safety-critical applications to prevent rare, yet catastrophic events. Unlike risk-neutral objectives, the static CVaR of the return depends on entire trajectories without admitting a recursive Bellman decomposition i…

Cited by 0SourceScholar
2025

Q-learning for Quantile MDPs: A Decomposition, Performance, and Convergence Analysis

AISTATS 2025poster

In Markov decision processes (MDPs), quantile risk measures such as Value-at-Risk are a standard metric for modeling RL agents' preferences for certain outcomes. This paper proposes a new Q-learning algorithm for quantile optimization in MDPs with strong convergence and performance guarantees. The a…

Cited by 0SourcecodeScholar
2025

State Entropy Regularization for Robust Reinforcement Learning

NeurIPS 2025oral

State entropy regularization has empirically shown better exploration and sample complexity in reinforcement learning (RL). However, its theoretical guarantees have not been studied. In this paper, we show that state entropy regularization improves robustness to structured and spatially correlated p…

Cited by 0SourceScholar
2024

Solving Non-rectangular Reward-Robust MDPs via Frequency Regularization

AAAI 2024technical

In robust Markov decision processes (RMDPs), it is assumed that the reward and the transition dynamics lie in a given uncertainty set. By targeting maximal return under the most adversarial model from that set, RMDPs address performance sensitivity to misspecified environments. Yet, to preserve comp…

Cited by 2SourcePDFScholar
2024

Tree Search-Based Policy Optimization under Stochastic Execution Delay

ICLR 2024poster

The standard formulation of Markov decision processes (MDPs) assumes that the agent's decisions are executed immediately. However, in numerous realistic applications such as robotics or healthcare, actions are performed with a delay whose value can even be stochastic. In this work, we introduce stoc…

2023

Policy Gradient for Rectangular Robust Markov Decision Processes

NeurIPS 2023poster

Policy gradient methods have become a standard for training reinforcement learning agents in a scalable and efficient manner. However, they do not account for transition uncertainty, whereas learning robust policies can be computationally expensive. In this paper, we introduce robust policy gradient…

Cited by 36SourcePDFScholar
2021

Acting in Delayed Environments with Non-Stationary Markov Policies

ICLR 2021poster

The standard Markov Decision Process (MDP) formulation hinges on the assumption that an action is executed immediately after it was chosen. However, assuming it is often unrealistic and can lead to catastrophic failures in applications such as robotic manipulation, cloud computing, and finance. We i…

2021

Twice regularized MDPs and the equivalence between robustness and regularization

NeurIPS 2021poster

Robust Markov decision processes (MDPs) aim to handle changing or partially known system dynamics. To solve them, one typically resorts to robust optimization methods. However, this significantly increases computational complexity and limits scalability in both learning and planning. On the other ha…

Cited by 51SourcePDFScholar