← Search

Khaled Eldowa

5 accepted papers

2025

Instance-Dependent Regret Bounds for Nonstochastic Linear Partial Monitoring

NeurIPS 2025poster

In contrast to the classic formulation of partial monitoring, linear partial monitoring can model infinite outcome spaces, while imposing a linear structure on both the losses and the observations. This setting can be viewed as a generalization of linear bandits where loss and feedback are decoupled…

Cited by 0SourceScholar
2025

Online Episodic Convex Reinforcement Learning

ICML 2025poster

We study online learning in episodic finite-horizon Markov decision processes (MDPs) with convex objective functions, known as the concave utility reinforcement learning (CURL) problem. This setting generalizes RL from linear to convex losses on the state-action distribution induced by the agent’s p…

Cited by 0SourcePDFScholar
2023

On the Minimax Regret for Online Learning with Feedback Graphs

NeurIPS 2023spotlight

In this work, we improve on the upper and lower bounds for the regret of online learning with strongly observable undirected feedback graphs. The best known upper bound for this problem is $\mathcal{O}\bigl(\sqrt{\alpha T\ln K}\bigr)$, where $K$ is the number of actions, $\alpha$ is the independence…

Cited by 11SourcePDFScholar
2022

Finite Sample Analysis of Mean-Volatility Actor-Critic for Risk-Averse Reinforcement Learning

AISTATS 2022poster

The goal in the standard reinforcement learning problem is to find a policy that optimizes the expected return. However, such an objective is not adequate in a lot of real-life applications, like finance, where controlling the uncertainty of the outcome is imperative. The mean-volatility objective p…