← Search

Xuzhao Shi

2 accepted papers

2025

QAEval: Mixture of Evaluators for Question-Answering Task Evaluation

ACL 2025long

Question answering (QA) tasks serve as a key benchmark for evaluating generation systems. Traditional rule-based metrics, such as accuracy and relaxed-accuracy, struggle with open-ended and unstructured responses. LLM-based evaluation methods offer greater flexibility but suffer from sensitivity to…

2024

SarcNet: A Multilingual Multimodal Sarcasm Detection Dataset

COLING 2024main

Sarcasm poses a challenge in linguistic analysis due to its implicit nature, involving an intended meaning that contradicts the literal expression. The advent of social networks has propelled the utilization of multimodal data to enhance sarcasm detection performance. In prior multimodal sarcasm det…