← Search

Stefan Stojanovic

5 accepted papers

2025

Shift Before You Learn: Enabling Low-Rank Representations in Reinforcement Learning

NeurIPS 2025spotlight

Low-rank structure is a common implicit assumption in many modern reinforcement learning (RL) algorithms. For instance, reward-free and goal-conditioned RL methods often presume that the successor measure admits a low-rank representation. In this work, we challenge this assumption by first remarking…

Cited by 0SourceScholar
2024

Low-Rank Bandits via Tight Two-to-Infinity Singular Subspace Recovery

ICML 2024poster

We study contextual bandits with low-rank structure where, in each round, if the (context, arm) pair $(i,j)\in [m]\times [n]$ is selected, the learner observes a noisy sample of the $(i,j)$-th entry of an unknown low-rank reward matrix. Successive contexts are generated randomly in an i.i.d. manner…

2024

Model-free Low-Rank Reinforcement Learning via Leveraged Entry-wise Matrix Estimation

NeurIPS 2024poster

We consider the problem of learning an $\varepsilon$-optimal policy in controlled dynamical systems with low-rank latent structure. For this problem, we present LoRa-PI (Low-Rank Policy Iteration), a model-free learning algorithm alternating between policy improvement and policy evaluation steps. I…

Cited by 1SourcePDFScholar
2023

Spectral Entry-wise Matrix Estimation for Low-Rank Reinforcement Learning

NeurIPS 2023poster

We study matrix estimation problems arising in reinforcement learning with low-rank structure. In low-rank bandits, the matrix to be recovered specifies the expected arm rewards, and for low-rank Markov Decision Processes (MDPs), it characterizes the transition kernel of the MDP. In both cases, each…

Cited by 11SourcePDFScholar
2022

Fast rates for noisy interpolation require rethinking the effect of inductive bias

ICML 2022spotlight

Good generalization performance on high-dimensional data crucially hinges on a simple structure of the ground truth and a corresponding strong inductive bias of the estimator. Even though this intuition is valid for regularized models, in this paper we caution against a strong inductive bias for int…

Cited by 26SourcePDFScholar