← Search

Timo Flesch

1 accepted papers

2026

Quantifying Biases in LLM-as-a-Judge Evaluations

ICML 2026poster

The evaluation of large language models (LLMs) is increasingly performed by other LLMs, a setup commonly known as "LLM-as-a-judge", or autograders. While autograders offer a scalable alternative to human evaluation, they are not free from biases (e.g., favouring longer outputs or generations from th…

Cited by 0SourceScholar