← Search

Jinglin Teh

1 accepted papers

2026

Breaking Safety Paradox with Feasible Dual Policy Iteration

ICLR 2026poster

Achieving zero constraint violations in safe reinforcement learning poses a significant challenge. We discover a key obstacle called the safety paradox, where improving policy safety reduces the frequency of constraint-violating samples, thereby impairing feasibility function estimation and ultimate…

Cited by 0SourceScholar