← Search

Julia Olkhovskaya

6 accepted papers

2025

An Improved Algorithm for Adversarial Linear Contextual Bandits via Reduction

NeurIPS 2025poster

We present an efficient algorithm for linear contextual bandits with adversarial losses and stochastic action sets. Our approach reduces this setting to misspecification-robust adversarial linear bandits with fixed action sets. Without knowledge of the context distribution or access to a context sim…

Cited by 0SourceScholar
2024

Kernel-Based Function Approximation for Average Reward Reinforcement Learning: An Optimist No-Regret Algorithm

NeurIPS 2024poster

Reinforcement Learning (RL) utilizing kernel ridge regression to predict the expected value function represents a powerful method with great representational capacity. This setting is a highly versatile framework amenable to analytical results. We consider kernel-based function approximation for RL…

Cited by 0SourcePDFScholar
2023

First- and Second-Order Bounds for Adversarial Linear Contextual Bandits

NeurIPS 2023poster

We consider the adversarial linear contextual bandit setting, which allows for the loss functions associated with each of $K$ arms to change over time without restriction. Assuming the $d$-dimensional contexts are drawn from a fixed known distribution, the worst-case expected regret over the course…

Cited by 10SourcePDFScholar
2022

Lifting the Information Ratio: An Information-Theoretic Analysis of Thompson Sampling for Contextual Bandits

NeurIPS 2022accept

We study the Bayesian regret of the renowned Thompson Sampling algorithm in contextual bandits with binary losses and adversarially-selected contexts. We adapt the information-theoretic perspective of Russo and Van Roy [2016] to the contextual setting by considering a lifted version of the informati…

Cited by 0SourcePDFScholar
2021

Online learning in MDPs with linear function approximation and bandit feedback.

NeurIPS 2021poster

We consider the problem of online learning in an episodic Markov decision process, where the reward function is allowed to change between episodes in an adversarial manner and the learner only observes the rewards associated with its actions. We assume that rewards and the transition function can be…

Cited by 34SourcePDFScholar