2025
Risk-aware Direct Preference Optimization under Nested Risk Measure
NeurIPS 2025poster
When fine-tuning pre-trained Large Language Models (LLMs) to align with human values and intentions, maximizing the estimated reward can lead to superior performance, but it also introduces potential risks due to deviations from the reference model's intended behavior. Most existing methods typicall…