← Search

Benjamin Genchel

1 accepted papers

2025

Language Model Council: Democratically Benchmarking Foundation Models on Highly Subjective Tasks

NAACL 2025long

As Large Language Models (LLMs) continue to evolve, evaluating them remains a persistent challenge. Many recent evaluations use LLMs as judges to score outputs from other LLMs, often relying on a single large model like GPT-4o. However, using a single LLM judge is prone to intra-model bias, and many…