← Search

Ronald Xu

1 accepted papers

2024

Efficient multi-prompt evaluation of LLMs

NeurIPS 2024poster

Most popular benchmarks for comparing LLMs rely on a limited set of prompt templates, which may not fully capture the LLMs’ abilities and can affect the reproducibility of results on leaderboards. Many recent works empirically verify prompt sensitivity and advocate for changes in LLM evaluation. In…

Cited by 12SourcePDFScholar