2026
Towards Achieving Optimal Strong Regret and Constraint Violation via Computational Efficient Model-free RL
ICML 2026poster
We study episodic constrained Markov decision processes (CMDPs) with linear function approximation, where the goal is to achieve strong regret and constraint violation guarantees without allowing error cancellations. Unlike the existing work, which focuses on either tabular CMDP or model-based reinf…