← Search

Ankan Saha

2 accepted papers

2025

AlphaPO: Reward Shape Matters for LLM Alignment

ICML 2025poster

Reinforcement Learning with Human Feedback (RLHF) and its variants have made huge strides toward the effective alignment of large language models (LLMs) to follow instructions and reflect human values. More recently, Direct Alignment Algorithms (DAAs) have emerged in which the reward modeling stage…

Cited by 0SourcePDFScholar
2017

Large-Scale Quadratically Constrained Quadratic Program via Low-Discrepancy Sequences

NeurIPS 2017poster

We consider the problem of solving a large-scale Quadratically Constrained Quadratic Program. Such problems occur naturally in many scientific and web applications. Although there are efficient methods which tackle this problem, they are mostly not scalable. In this paper, we develop a method that t…

Cited by 11SourcePDFScholar