← Search

Peyman Mohajerin Esfahani

10 accepted papers

2026

G$^2$RPO: Geometric GRPO; Escaping LLM's Reasoning Rut to Break Accuracy--Entropy Trade-off

ICML 2026poster

Reinforcement learning with verifiable rewards (RLVR) is a cornerstone of post-training for large reasoning models, yet widely used algorithms such as Group Relative Policy Optimization (GRPO) often exhibit \textbf{diversity collapse}. We provide a geometric diagnosis by formalizing GRPO as a dynami…

Cited by 0SourceScholar
2026

Rate or Fate? RLV$^{\varepsilon}$R: Reinforcement Learning with Verifiable Noisy Rewards

ICML 2026spotlight

Reinforcement learning with verifiable rewards (RLVR) trains a policy by verifying sampled completions and reinforcing higher-scoring outputs, but practical verifiers (e.g., incomplete unit tests or noisy judges) are prone to false positives and false negatives. We ask when such noise merely slows l…

Cited by 0SourceScholar
2025

Rank-One Modified Value Iteration

ICML 2025poster

In this paper, we provide a novel algorithm for solving planning and learning problems of Markov decision processes. The proposed algorithm follows a policy iteration-type update by using a rank-one approximation of the transition probability matrix in the policy evaluation step. This rank-one app…

Cited by 0SourcePDFScholar
2024

Scalable Kernel Inverse Optimization

NeurIPS 2024poster

Inverse Optimization (IO) is a framework for learning the unknown objective function of an expert decision-maker from a past dataset. In this paper, we extend the hypothesis class of IO objective functions to a reproducing kernel Hilbert space (RKHS), thereby enhancing feature representation to an i…

2021

Fast Approximate Dynamic Programming for Infinite-Horizon Markov Decision Processes

NeurIPS 2021poster

In this study, we consider the infinite-horizon, discounted cost, optimal control of stochastic nonlinear systems with separable cost and constraints in the state and input variables. Using the linear-time Legendre transform, we propose a novel numerical scheme for implementation of the correspondin…

2021

Principal Component Hierarchy for Sparse Quadratic Programs

ICML 2021spotlight

We propose a novel approximation hierarchy for cardinality-constrained, convex quadratic programs that exploits the rank-dominating eigenvectors of the quadratic matrix. Each level of approximation admits a min-max characterization whose objective function can be optimized over the binary variables…

2018

Fast Gradient-Based Methods with Exponential Rate: A Hybrid Control Framework

ICML 2018oral

Ordinary differential equations, and in general a dynamical system viewpoint, have seen a resurgence of interest in developing fast optimization methods, mainly thanks to the availability of well-established analysis tools. In this study, we pursue a similar objective and propose a class of hybrid c…

Cited by 16SourcePDFScholar
2018

Wasserstein Distributionally Robust Kalman Filtering

NeurIPS 2018spotlight

We study a distributionally robust mean square error estimation problem over a nonconvex Wasserstein ambiguity set containing only normal distributions. We show that the optimal estimator and the least favorable distribution form a Nash equilibrium. Despite the non-convex nature of the ambiguity set…

2015

Distributionally Robust Logistic Regression

NeurIPS 2015spotlight

This paper proposes a distributionally robust approach to logistic regression. We use the Wasserstein distance to construct a ball in the space of probability distributions centered at the uniform distribution on the training samples. If the radius of this Wasserstein ball is chosen judiciously, we…

Cited by 401SourcePDFScholar