← Search

Grace Byun

2 accepted papers

2025

D-GEN: Automatic Distractor Generation and Evaluation for Reliable Assessment of Generative Models

ACL 2025finding

Evaluating generative models with open-ended generation is challenging due to inconsistencies in response formats. Multiple-choice (MC) evaluation mitigates this issue, but generating high-quality distractors is time-consuming and labor-intensive. We introduce D-GEN, the first open-source distractor…

2025

Measuring Sycophancy of Language Models in Multi-turn Dialogues

EMNLP 2025

Large Language Models (LLMs) are expected to provide helpful and harmless responses, yet they often exhibit sycophancy —conforming to user beliefs regardless of factual accuracy or ethical soundness. Prior research on sycophancy has primarily focused on single-turn factual correctness, overlooking t