← Search

Alon Peled-Cohen

4 accepted papers

2026

Near-Optimal Regret for Policy Optimization in Contextual MDPs with General Offline Function Approximation

ICML 2026poster

We introduce OPO-CMDP, the first policy optimization algorithm for stochastic Contextual Markov Decision Process (CMDPs) under general offline function approximation. Our approach achieves a high probability regret bound of $\widetilde{O}(H^4\sqrt{T|S||A|\log(|\mathcal{F}||\mathcal{P}|)}),$ where $S…

Cited by 0SourceScholar
2021

Learning and Evaluating a Differentially Private Pre-trained Language Model

EMNLP 2021finding

Contextual language models have led to significantly better results, especially when pre-trained on the same data as the downstream task. While this additional pre-training usually improves performance, it can lead to information leakage and therefore risks the privacy of individuals mentioned in th…

Cited by 78SourcePDFScholar