UAI 2024poster0 citations

Offline Reward Perturbation Boosts Distributional Shift in Online RL

Zishun Yu, Siteng Kang, Xinhua Zhang

Abstract

Offline-to-online reinforcement learning has recently been shown effective in reducing the online sample complexity by first training from offline collected data. However, this additional data source may also invite new poisoning attacks that target offline training. In this work, we reveal such vulnerabilities in

BibTeX
@InProceedings{pmlr-v244-yu24a,
  title = 	 {Offline Reward Perturbation Boosts Distributional Shift in Online RL},
  author =       {Yu, Zishun and Kang, Siteng and Zhang, Xinhua},
  booktitle = 	 {Proceedings of the Fortieth Conference on Uncertainty in Artificial Intelligence},
  pages = 	 {4041--4055},
  year = 	 {2024},
  editor = 	 {Kiyavash, Negar and Mooij, Joris M.},
  volume = 	 {244},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {15--19 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://raw.githubusercontent.com/mlresearch/v244/main/assets/yu24a/yu24a.pdf},
  url = 	 {https://proceedings.mlr.press/v244/yu24a.html},
  abstract = 	 {Offline-to-online reinforcement learning has recently been shown effective in reducing the online sample complexity by first training from offline collected data. However, this additional data source may also invite new poisoning attacks that target offline training. In this work, we reveal such vulnerabilities in
Offline Reward Perturbation Boosts Distributional Shift in Online RL · UAI 2024