← Search

Raoyuan Zhao

4 accepted papers

2026

Human Uncertainty-Aware Data Selection and Automatic Labeling in Visual Question Answering

ICLR 2026poster

Large vision-language models (VLMs) achieve strong performance in Visual Question Answering but still rely heavily on supervised fine-tuning (SFT) with massive labeled datasets, which is costly due to human annotations. Crucially, real-world datasets often exhibit *human uncertainty* (**HU**) — var…

Cited by 0SourceScholar
2025

MAKIEval: A Multilingual Automatic WiKidata-based Framework for Cultural Awareness Evaluation for LLMs

EMNLP 2025

Large language models (LLMs) are used globally across many languages, but their English-centric pretraining raises concerns about cross-lingual disparities for cultural awareness, often resulting in biased outputs. However, comprehensive multilingual evaluation remains challenging due to limited ben

2025

What’s the Difference? Supporting Users in Identifying the Effects of Prompt and Model Changes Through Token Patterns

ACL 2025long

Prompt engineering for large language models is challenging, as even small prompt perturbations or model changes can significantly impact the generated output texts. Existing evaluation methods of LLM outputs, either automated metrics or human evaluation, have limitations, such as providing limited…

2024

SynthEval: Hybrid Behavioral Testing of NLP Models with Synthetic CheckLists

EMNLP 2024finding

Traditional benchmarking in NLP typically involves using static, held-out test sets and calculating aggregated statistics based on diverse examples. However, this approach often results in an overestimation of performance and lacks the ability to offer comprehensive, interpretable, and dynamic asses…