← Search

Taiye Chen

3 accepted papers

2025

Mitigating Reward Over-Optimization in RLHF via Behavior-Supported Regularization

ICLR 2025poster

Reinforcement learning from human feedback (RLHF) is an effective method for aligning large language models (LLMs) with human values. However, reward over-optimization remains an open challenge leading to discrepancies between the performance of LLMs under the reward model and the true human objecti…

Cited by 0SourcePDFScholar
2024

SafeSora: Towards Safety Alignment of Text2Video Generation via a Human Preference Dataset

NeurIPS 2024poster

To mitigate the risk of harmful outputs from large vision models (LVMs), we introduce the *SafeSora* dataset to promote research on aligning text-to-video generation with human values. This dataset encompasses human preferences in text-to-video generation tasks along two primary dimensions: helpfuln…