← Search

Molly Lin

1 accepted papers

2024

Rule Based Rewards for Language Model Safety

NeurIPS 2024poster

Reinforcement learning based fine-tuning of large language models (LLMs) on human preferences has been shown to enhance both their capabilities and safety behavior. However, in cases related to safety, without precise instructions to human annotators, the data collected may cause the model to beco…