2025
Amulet: Putting Complex Multi-Turn Conversations on the Stand with LLM Juries
EMNLP 2025
Today, large language models are widely used as judges to evaluate responses from other language models. Hence, it is imperative to benchmark and improve these LLM-judges on real-world language model usage: a typical human-assistant conversation is lengthy, and shows significant diversity in topics,