← Search

Andong Hua

3 accepted papers

2025

Flaw or Artifact? Rethinking Prompt Sensitivity in Evaluating LLMs

EMNLP 2025

Prompt sensitivity, referring to the phenomenon where paraphrasing (that is, repeating something written or spoken using different words) leads to significant changes in large language model performance, has been widely accepted as a core limitation of large language models. In this work, we revisit

Cited by 0SourcePDFScholar
2025

NutriBench: A Dataset for Evaluating Large Language Models in Nutrition Estimation from Meal Descriptions

ICLR 2025poster

Accurate nutrition estimation helps people make informed dietary choices and is essential in the prevention of serious health complications. We present NutriBench, the first publicly available natural language meal description nutrition benchmark. NutriBench consists of 11,857 meal descriptions gene…

2024

Initialization Matters for Adversarial Transfer Learning

CVPR 2024poster

With the prevalence of the Pretraining-Finetuning paradigm in transfer learning the robustness of downstream tasks has become a critical concern. In this work we delve into adversarial robustness in transfer learning and reveal the critical role of initialization including both the pretrained model…