← Search

Jixiang Hong

1 accepted papers

2024

CycleAlign: Iterative Distillation from Black-box LLM to White-box Models for Better Human Alignment

ACL 2024findings

Language models trained on large-scale corpus often generate harmful responses that are harmful and contrary to human values. A prevalent approach for human alignment is reinforcement learning from human feedback (RLHF), utilizing algorithms such as proximal policy optimization (PPO). However, these…