← Search

Shibo Chen

3 accepted papers

2026

GAS: Enhancing Reward-Cost Balance of Generative Model-assisted Offline Safe RL

ICLR 2026poster

Offline Safe Reinforcement Learning (OSRL) aims to learn a policy that achieves high performance in sequential decision-making while satisfying safety constraints, using only pre-collected datasets. Recent works, inspired by the strong capabilities of Generative Models (GMs), reformulate decision-ma…

Cited by 0SourceScholar
2025

Reinforcement Learning with Intrinsically Motivated Feedback Graph for Lost-sales Inventory Control

AISTATS 2025poster

Reinforcement learning (RL) has proven to be well-performed and versatile in inventory control (IC). However, further improvement of RL algorithms in the IC domain is impeded by two limitations of online experience. First, online experience is expensive to acquire in real-world applications. With th…

Cited by 0SourcecodeScholar
2024

Individual Contributions as Intrinsic Exploration Scaffolds for Multi-agent Reinforcement Learning

ICML 2024poster

In multi-agent reinforcement learning (MARL), effective exploration is critical, especially in sparse reward environments. Although introducing global intrinsic rewards can foster exploration in such settings, it often complicates credit assignment among agents. To address this difficulty, we propos…