← Search

Seyedali Mirjalili

1 accepted papers

2026

TUR-DPO: Topology- and Uncertainty-Aware Direct Preference Optimization

ICML 2026poster

Aligning large language models (LLMs) with human preferences is commonly done via reinforcement learning from human feedback (RLHF) with Proximal Policy Optimization (PPO) or, more simply, via Direct Preference Optimization (DPO). While DPO is stable and RL-free, it treats preferences as flat winner…

Cited by 0SourceScholar
Seyedali Mirjalili — accepted AI-conference papers · AIConfPaper