← Search

Shuyi Xie

2 accepted papers

2025

Diagnosing Failures in Large Language Models’ Answers: Integrating Error Attribution into Evaluation Framework

ACL 2025finding

With the widespread application of Large Language Models (LLMs) in various tasks, the mainstream LLM platforms generate massive user-model interactions daily. In order to efficiently analyze the performance of models and diagnose failures in their answers, it is essential to develop an automated fra…

2024

IDGen: Item Discrimination Induced Prompt Generation for LLM Evaluation

NeurIPS 2024poster

As Large Language Models (LLMs) become more capable of handling increasingly complex tasks, the evaluation set must keep pace with these advancements to ensure it remains sufficiently discriminative. Item Discrimination (ID) theory, which is widely used in educational assessment, measures the abilit…