← Search

Stephanie Chan

2 accepted papers

2026

Uncovering Competency Gaps in Large Language Models and Their Benchmarks

ICML 2026poster

The evaluation of large language models relies heavily on standardized benchmarks. These benchmarks provide useful aggregated metrics, but can obscure (i) particular sub-areas where the models are weak ("model gaps") (ii) imbalanced coverage in the benchmarks themselves ("benchmark gaps"). To automa…

Cited by 0SourceScholar
2022

Can language models learn from explanations in context?

EMNLP 2022finding

Language Models (LMs) can perform new tasks by adapting to a few in-context examples. For humans, explanations that connect examples to task principles can improve learning. We therefore investigate whether explanations of few-shot examples can help LMs. We annotate questions from 40 challenging tas…

Cited by 301SourcePDFScholar