← Search

Shahar Levy

4 accepted papers

2025

More Documents, Same Length: Isolating the Challenge of Multiple Documents in RAG

EMNLP 2025

Retrieval-Augmented Generation (RAG) enhances the accuracy of Large Language Model (LLM) responses by leveraging relevant external documents during generation. Although previous studies noted that retrieving many documents can degrade performance, they did not isolate how the quantity of documents a

2025

ReliableEval: A Recipe for Stochastic LLM Evaluation via Method of Moments

EMNLP 2025

LLMs are highly sensitive to prompt phrasing, yet standard benchmarks typically report performance using a single prompt, raising concerns about the reliability of such evaluations. In this work, we argue for a stochastic method of moments evaluation over the space of meaning-preserving prompt pertu

2021

Collecting a Large-Scale Gender Bias Dataset for Coreference Resolution and Machine Translation

EMNLP 2021finding

Recent works have found evidence of gender bias in models of machine translation and coreference resolution using mostly synthetic diagnostic datasets. While these quantify bias in a controlled experiment, they often do so on a small scale and consist mostly of artificial, out-of-distribution senten…

2021

Learning Disentangled Behavior Embeddings

NeurIPS 2021spotlight

To understand the relationship between behavior and neural activity, experiments in neuroscience often include an animal performing a repeated behavior such as a motor task. Recent progress in computer vision and deep learning has shown great potential in the automated analysis of behavior by levera…