2026
AutoMetrics: Approximate Human Judgments with Automatically Generated Evaluators
ICLR 2026poster
Evaluating user-facing AI applications remains a central challenge, especially in open-ended domains such as travel planning, clinical note generation, or dialogue. The gold standard is user feedback (e.g., thumbs up/down) or behavioral signals (e.g., retention), but these are often scarce in protot…