← Search

Sohhyung Park

2 accepted papers

2026

Diagnosing the Reliability of LLM-as-a-Judge via Item Response Theory

ICML 2026poster

While LLM-as-a-Judge is widely used in automated evaluation, existing validation practices primarily operate at the level of observed outputs, offering limited insight into whether LLM judges themselves function as stable and reliable measurement instruments. To address this limitation, we introduce…

Cited by 0SourceScholar
2025

LLM-guided Plan and Retrieval: A Strategic Alignment for Interpretable User Satisfaction Estimation in Dialogue

NAACL 2025long

Understanding user satisfaction with conversational systems, known as User Satisfaction Estimation (USE), is essential for assessing dialogue quality and enhancing user experiences. However, existing methods for USE face challenges due to limited understanding of underlying reasons for user dissatis…

Cited by 0SourcePDFScholar