← Search

Yaacov Ritov

2 accepted papers

2025

A transfer learning framework for weak to strong generalization

ICLR 2025poster

Modern large language model (LLM) alignment techniques rely on human feedback, but it is unclear whether the techniques fundamentally limit the capabilities of aligned LLMs. In particular, it is unclear whether it is possible to align (stronger) LLMs with superhuman capabilities with (weaker) human…

Cited by 0SourcePDFScholar
2024

Learning in reverse causal strategic environments with ramifications on two sided markets

ICLR 2024poster

Motivated by equilibrium models of labor markets, we develop a formulation of causal strategic classification in which strategic agents can directly manipulate their outcomes. As an application, we consider employers that seek to anticipate the strategic response of a labor force when developing a h…

Cited by 3SourcePDFScholar