← Search

Harvey Yiyun Fu

3 accepted papers

2025

Absence Bench: Language Models Can’t See What’s Missing

NeurIPS 2025spotlight

Large language models (LLMs) are increasingly capable of processing long inputs and locating specific information within them, as evidenced by their performance on the Needle in a Haystack (NIAH) test. However, while models excel at recalling surprising information, they still struggle to identify c…

Cited by 0SourceScholar
2023

Estimating Large Language Model Capabilities without Labeled Test Data

EMNLP 2023long findings

Large Language Models (LLMs) have exhibited an impressive ability to perform in-context learning (ICL) from only a few examples, but the success of ICL varies widely from task to task. Thus, it is important to quickly determine whether ICL is applicable to a new task, but directly evaluating ICL acc…

Cited by 0SourcecodeScholar
2023

How Predictable Are Large Language Model Capabilities? A Case Study on BIG-bench

EMNLP 2023long findings

We investigate the predictability of large language model (LLM) capabilities: given records of past experiments using different model families, numbers of parameters, tasks, and numbers of in-context examples, can we accurately predict LLM performance on new experiment configurations? Answering this…

Cited by 0SourcecodeScholar