← Search

Vivek Datla

1 accepted papers

2026

Alignment-Weighted DPO: A principled reasoning approach to improve alignment

ICLR 2026poster

Recent advances in alignment techniques such as Supervised Fine-Tuning (SFT), Reinforcement Learning from Human Feedback (RLHF), and Direct Preference Optimization (DPO) have improved the safety of large language models (LLMs). However, these LLMs remain vulnerable to jailbreak attacks that disguise…

Cited by 0SourceScholar