← Search

Timo Kaufmann

5 accepted papers

2026

Calibrated Preference Learning: The Case of Label Ranking

ICML 2026poster

Calibration, the alignment of predicted probabilities with true outcome frequencies, is essential for reliable decision-making. While extensively studied for classification and regression, calibration has not been formally addressed for probabilistic label ranking, where the goal is to predict a dis…

Cited by 0SourceScholar
2025

Comparing Comparisons: Informative and Easy Human Feedback with Distinguishability Queries

ICML 2025poster

Learning human objectives from preference feedback has significantly advanced reinforcement learning (RL) in domains where objectives are hard to formalize. However, traditional methods based on pairwise trajectory comparisons face notable challenges, including the difficulty in comparing trajector…

Cited by 1SourcePDFScholar
2025

DUO: Diverse, Uncertain, On-Policy Query Generation and Selection for Reinforcement Learning from Human Feedback

AAAI 2025technical

Defining a reward function is usually a challenging but critical task for the system designer in reinforcement learning, especially when specifying complex behaviors. Reinforcement learning from human feedback (RLHF) emerges as a promising approach to circumvent this. In RLHF, the agent typically le…

Cited by 0SourcePDFScholar
2025

Inverse Constitutional AI: Compressing Preferences into Principles

ICLR 2025poster

Feedback data is widely used for fine-tuning and evaluating state-of-the-art AI models. Pairwise text preferences, where human or AI annotators select the “better” of two options, are particularly common. Such preferences are used to train (reward) models or to rank models with aggregate statistics.…

2025

ResponseRank: Data-Efficient Reward Modeling through Preference Strength Learning

NeurIPS 2025poster

Binary choices, as often used for reinforcement learning from human feedback (RLHF), convey only the *direction* of a preference. A person may choose apples over oranges and bananas over grapes, but *which preference is stronger*? Strength is crucial for decision-making under uncertainty and general…

Cited by 0SourceScholar