← Search

Junhyuk Choi

3 accepted papers

2026

Diagnosing the Reliability of LLM-as-a-Judge via Item Response Theory

ICML 2026poster

While LLM-as-a-Judge is widely used in automated evaluation, existing validation practices primarily operate at the level of observed outputs, offering limited insight into whether LLM judges themselves function as stable and reliable measurement instruments. To address this limitation, we introduce…

Cited by 0SourceScholar
2025

People will agree what I think: Investigating LLM’s False Consensus Effect

NAACL 2025findings

Large Language Models (LLMs) have been recently adopted in interactive systems requiring communication. As the false belief in a model can harm the usability of such systems, LLMs should not have cognitive biases that humans have. Psychologists especially focus on the False Consensus Effect (FCE), a…

2025

VoiceBBQ: Investigating Effect of Content and Acoustics in Social Bias of Spoken Language Model

EMNLP 2025

We introduce VoiceBBQ, a spoken extension of the BBQ (Bias Benchmark for Question answering) - a dataset that measures social bias by presenting ambiguous or disambiguated contexts followed by questions that may elicit stereotypical responses. Due to the nature of speech modality, social bias in Spo