← Search

Saloni Mittal

1 accepted papers

2025

Multi-Reference Preference Optimization for Large Language Models

AAAI 2025technical

How can Large Language Models (LLMs) be aligned with human intentions and values? A typical solution is to gather human preference on model outputs and finetune the LLMs accordingly while ensuring that updates do not deviate too far from a reference model. Recent approaches, such as direct preferenc…