← Search

Philip Wang

1 accepted papers

2026

Belief-Based Offline Reinforcement Learning for Delay-Robust Policy Optimization

ICLR 2026poster

Offline–to–online deployment of reinforcement learning (RL) agents often stumbles over two fundamental gaps: (1) the sim-to-real gap, where real-world systems exhibit latency and other physical imperfections not captured in simulation; and (2) the interaction gap, where policies trained purely offli…

Cited by 0SourcecodeScholar