← Search

Emmanuel Esposito

6 accepted papers

2026

Parameter-free Dynamic Regret: Time-varying Movement Costs, Delayed Feedback, and Memory

ICML 2026poster

In this paper, we study dynamic regret in unconstrained online convex optimization (OCO) with movement costs. Specifically, we generalize the standard setting by allowing the movement cost coefficients $\lambda_t$ to vary arbitrarily over time. Our main contribution is a novel algorithm that establi…

Cited by 0SourceScholar
2025

Exploiting Curvature in Online Convex Optimization with Delayed Feedback

ICML 2025poster

In this work, we study the online convex optimization problem with curved losses and delayed feedback. When losses are strongly convex, existing approaches obtain regret bounds of order $d_{\max} \ln T$, where $d_{\max}$ is the maximum delay and $T$ is the time horizon. However, in many cases, this…

Cited by 0SourcePDFScholar
2025

When Lower-Order Terms Dominate: Adaptive Expert Algorithms for Heavy-Tailed Losses

NeurIPS 2025poster

We consider the problem setting of prediction with expert advice with possibly heavy-tailed losses, i.e.\ the only assumption on the losses is an upper bound on their second moments, denoted by $\theta$. We develop adaptive algorithms that do not require any prior knowledge about the range or the se…

Cited by 0SourceScholar
2023

Delayed Bandits: When Do Intermediate Observations Help?

ICML 2023poster

We study a $K$-armed bandit with delayed feedback and intermediate observations. We consider a model, where intermediate observations have a form of a finite state, which is observed immediately after taking an action, whereas the loss is observed after an adversarially chosen delay. We show that th…

Cited by 2SourcePDFScholar
2023

On the Minimax Regret for Online Learning with Feedback Graphs

NeurIPS 2023spotlight

In this work, we improve on the upper and lower bounds for the regret of online learning with strongly observable undirected feedback graphs. The best known upper bound for this problem is $\mathcal{O}\bigl(\sqrt{\alpha T\ln K}\bigr)$, where $K$ is the number of actions, $\alpha$ is the independence…

Cited by 11SourcePDFScholar
2022

Learning on the Edge: Online Learning with Stochastic Feedback Graphs

NeurIPS 2022accept

The framework of feedback graphs is a generalization of sequential decision-making with bandit or full information feedback. In this work, we study an extension where the directed feedback graph is stochastic, following a distribution similar to the classical Erdős-Rényi model. Specifically, in each…

Cited by 16SourcePDFScholar