← Search

Cristina Garbacea

4 accepted papers

2025

RATE: Causal Explainability of Reward Models with Imperfect Counterfactuals

ICML 2025poster

Reward models are widely used as proxies for human preferences when aligning or evaluating LLMs. However, reward models are black boxes, and it is often unclear what, exactly, they are actually rewarding. In this paper we develop Rewrite-based Attribute Treatment Estimator (RATE) as an effective met…

2024

BoNBoN Alignment for Large Language Models and the Sweetness of Best-of-n Sampling

NeurIPS 2024poster

This paper concerns the problem of aligning samples from large language models to human preferences using *best-of-$n$* sampling, where we draw $n$ samples, rank them, and return the best one. We consider two fundamental problems. First: what is the relationship between best-of-$n$ and other (RLHF-t…

Cited by 28SourcePDFScholar
2021

Explainable Prediction of Text Complexity: The Missing Preliminaries for Text Simplification

ACL 2021long

Text simplification reduces the language complexity of professional content for accessibility purposes. End-to-end neural network models have been widely adopted to directly generate the simplified version of input text, usually functioning as a blackbox. We show that text simplification can be deco…

Cited by 30SourcePDFScholar