← Search

Arduin Findeis

2 accepted papers

2025

Can External Validation Tools Improve Annotation Quality for LLM-as-a-Judge?

ACL 2025long

Pairwise preferences over model responses are widely collected to evaluate and provide feedback to large language models (LLMs). Given two alternative model responses to the same input, a human or AI annotator selects the “better” response. This approach can provide feedback for domains where other…

2025

Inverse Constitutional AI: Compressing Preferences into Principles

ICLR 2025poster

Feedback data is widely used for fine-tuning and evaluating state-of-the-art AI models. Pairwise text preferences, where human or AI annotators select the “better” of two options, are particularly common. Such preferences are used to train (reward) models or to rank models with aggregate statistics.…