← Search

Yotam Wolf

1 accepted papers

2024

Fundamental Limitations of Alignment in Large Language Models

ICML 2024poster

An important aspect in developing language models that interact with humans is aligning their behavior to be useful and unharmful for their human users. This is usually achieved by tuning the model in a way that enhances desired behaviors and inhibits undesired ones, a process referred to as alignme…