← Search

Frank Yang

2 accepted papers

2026

Belief-Based Offline Reinforcement Learning for Delay-Robust Policy Optimization

ICLR 2026poster

Offline–to–online deployment of reinforcement learning (RL) agents often stumbles over two fundamental gaps: (1) the sim-to-real gap, where real-world systems exhibit latency and other physical imperfections not captured in simulation; and (2) the interaction gap, where policies trained purely offli…

Cited by 0SourcecodeScholar
2026

Delayed Feedback Modeling with Influence Functions

AAAI 2026technical

In online advertising under the cost-per-conversion (CPA) model, accurate conversion rate (CVR) prediction is crucial. A major challenge is delayed feedback, where conversions may occur long after user interactions, leading to incomplete recent data and biased model training. Existing solutions part

Cited by 0SourcePDFScholar