← Search

Mohit Raghavendra

1 accepted papers

2025

Balancing the Budget: Understanding Trade-offs Between Supervised and Preference-Based Finetuning

ACL 2025long

Post-training of Large Language Models often involves a pipeline of Supervised Finetuning (SFT) followed by Preference Finetuning (PFT) using methods like Direct Preference Optimization. Both stages require annotated data that are very different in structure and costs. We study how to optimally allo…