A Unified Theoretical Analysis of Private and Robust Offline Alignment: from RLHF to DPO
In this paper, we theoretically investigate the effects of noisy labels in offline alignment, with a focus on the interplay between privacy and robustness against adversarial corruption. Specifically, under linear modeling assumptions, we present a unified analysis covering both reinforcement learni…