← Search

Idan Pipano

3 accepted papers

2026

Displacement-Resistant Extensions of DPO with Nonconvex $f$-Divergences

ICLR 2026poster

DPO and related algorithms align language models by directly optimizing the RLHF objective: find a policy that maximizes the Bradley-Terry reward while staying close to a reference policy through a KL divergence penalty. Previous work showed that this approach could be further generalized: the origi…

Cited by 0SourceScholar
2026

The Search for Stability: Learning Dynamics of Strategic Publishers with Initial Documents (Abstract Reprint)

AAAI 2026technical

We study a game-theoretic information retrieval model in which strategic publishers aim to maximize their chances of being ranked first by the search engine while maintaining the integrity of their original documents. We show that the commonly used Probability Ranking Principle (PRP) ranking scheme

Cited by 0SourcePDFScholar
2025

On the Convergence of No-Regret Dynamics in Information Retrieval Games with Proportional Ranking Functions

ICLR 2025poster

Publishers who publish their content on the web act strategically, in a behavior that can be modeled within the online learning framework. Regret, a central concept in machine learning, serves as a canonical measure for assessing the performance of learning agents within this framework. We prove th…

Cited by 2SourcePDFScholar