← Search

Prashanth La

10 accepted papers

2024

A Cubic-regularized Policy Newton Algorithm for Reinforcement Learning

AISTATS 2024poster

We consider the problem of control in the setting of reinforcement learning (RL), where model information is not available. Policy gradient algorithms are a popular solution approach for this problem and are usually shown to converge to a stationary point of the value function. In this paper, we pro…

Cited by 3SourcePDFScholar
2023

Finite time analysis of temporal difference learning with linear function approximation: Tail averaging and regularisation

AISTATS 2023poster

We study the finite-time behaviour of the popular temporal difference (TD) learning algorithm, when combined with tail-averaging. We derive finite time bounds on the parameter error of the tail-averaged TD iterate under a step-size choice that does not require information about the eigenvalues of th…

Cited by 25SourcePDFScholar
2020

Concentration bounds for CVaR estimation: The cases of light-tailed and heavy-tailed distributions

ICML 2020poster

Conditional Value-at-Risk (CVaR) is a widely used risk metric in applications such as finance. We derive concentration bounds for CVaR estimates, considering separately the cases of sub-Gaussian, light-tailed and heavy-tailed distributions. For the sub-Gaussian and light-tailed cases, we use a class…

Cited by 70SourcePDFScholar
2016

(Bandit) Convex Optimization with Biased Noisy Gradient Oracles

AISTATS 2016poster

A popular class of algorithms for convex optimization and online learning with bandit feedback rely on constructing noisy gradient estimates, which are then used in place of the actual gradients in appropriately adjusted first-order algorithms. Depending on the properties of the function to be optim…

Cited by 18SourcePDFScholar
2016

Cumulative Prospect Theory Meets Reinforcement Learning: Prediction and Control

ICML 2016poster

Cumulative prospect theory (CPT) is known to model human decisions well, with substantial empirical evidence supporting this claim. CPT works by distorting probabilities and is more general than the classic expected utility and coherent risk measures. We bring this idea to a risk-sensitive reinforce…

Cited by 102SourcePDFScholar
2015

On TD(0) with function approximation: Concentration bounds and a centered variant with exponential convergence

ICML 2015poster

We provide non-asymptotic bounds for the well-known temporal difference learning algorithm TD(0) with linear function approximators. These include high-probability bounds as well as bounds in expectation. Our analysis suggests that a step-size inversely proportional to the number of iterations canno…

Cited by 61SourcePDFScholar