← Search

Yassir Jedra

9 accepted papers

2026

Two-Layer Linear Auto-Regressive Models Estimate Latent States

ICML 2026poster

Auto-regressive models have emerged as powerful tools for sequential data, from language to video. Understanding how and why these models learn latent representations remains an open theoretical question. In this work, we demonstrate that when trained by empirical risk minimization on data from part…

Cited by 0SourceScholar
2025

Optimal Transfer Learning for Missing Not-at-Random Matrix Completion

ICML 2025poster

We study transfer learning for matrix completion in a Missing Not-at-Random (MNAR) setting that is motivated by biological problems. The target matrix $Q$ has entire rows and columns missing, making estimation impossible without side information. To address this, we use a noisy and incomplete sourc…

Cited by 0SourcePDFScholar
2024

Low-Rank Bandits via Tight Two-to-Infinity Singular Subspace Recovery

ICML 2024poster

We study contextual bandits with low-rank structure where, in each round, if the (context, arm) pair $(i,j)\in [m]\times [n]$ is selected, the learner observes a noisy sample of the $(i,j)$-th entry of an unknown low-rank reward matrix. Successive contexts are generated randomly in an i.i.d. manner…

2024

Model-free Low-Rank Reinforcement Learning via Leveraged Entry-wise Matrix Estimation

NeurIPS 2024poster

We consider the problem of learning an $\varepsilon$-optimal policy in controlled dynamical systems with low-rank latent structure. For this problem, we present LoRa-PI (Low-Rank Policy Iteration), a model-free learning algorithm alternating between policy improvement and policy evaluation steps. I…

Cited by 1SourcePDFScholar
2023

Nearly Optimal Latent State Decoding in Block MDPs

AISTATS 2023poster

We consider the problem of model estimation in episodic Block MDPs. In these MDPs, the decision maker has access to rich observations or contexts generated from a small number of latent states. We are interested in estimating the latent state decoding function (the mapping from the observations to l…

2023

Spectral Entry-wise Matrix Estimation for Low-Rank Reinforcement Learning

NeurIPS 2023poster

We study matrix estimation problems arising in reinforcement learning with low-rank structure. In low-rank bandits, the matrix to be recovered specifies the expected arm rewards, and for low-rank Markov Decision Processes (MDPs), it characterizes the transition kernel of the MDP. In both cases, each…

Cited by 11SourcePDFScholar
2020

Optimal Algorithms for Multiplayer Multi-Armed Bandits

AISTATS 2020poster

The paper addresses various Multiplayer Multi-Armed Bandit (MMAB) problems, where M decision-makers, or players, collaborate to maximize their cumulative reward. We first investigate the MMAB problem where players selecting the same arms experience a collision (and are aware of it) and do not collec…

Cited by 93SourcePDFScholar