2025
RoDEval: A Robust Word Sense Disambiguation Evaluation Framework for Large Language Models
EMNLP 2025
Accurately evaluating the word sense disambiguation (WSD) capabilities of large language models (LLMs) remains challenging, as existing studies primarily rely on single-task evaluations and classification-based metrics that overlook the fundamental differences between generative LLMs and traditional